-
"A benchmark you really don't want models to be saturated with." Counts the number of times when AI agents inadvertently compromise or affect third-party entities
Tags: felonies hacking crime frontier-models llms anthropic openai meta evaluation testing benchmarks funny