The marginal value of another AI "safety" eval has collapsed. Of 185 published findings about a named AI company, one led to binding policy action: it was the narrow jailbreak flagged by Amazon ("fix this code") that got the US gov to pull Fable for 3 weeks.*
(* from a poster I saw at MATS by Kunal Singh (@kunalssingh_) and @StephenLCasper, who traced 1,002 published eval findings, 185 of them about a named company, see below)
We now get many more warning signs from incidents and capability marketing than from evals. From the last few months:
- INCIDENT: One guy, using AI agents, likely hacked at least nine South Korean banks in about two weeks.
- CAPABILITY: OpenAI just published full or partial solutions to 370+ open maths problems, at about 3 hours of compute on average.
- CAPABILITY: Claude now leads 26% of Anthropic's AI R&D work, up from almost nothing at the start of the year.
- EVAL: Biorisk evals are saturated, and the next level would be making actually dangerous things in wet labs.
- CAPABILITY: In August, a Stanford team used an AI model to design 16 new viable bacteriophages. Published in Science.
- INCIDENT: Empirical analysis after Hugging Face, shows that loss of control is fairly plausible. OpenAI's own agents also posted 53 ChatGPT users' images online.
- EVAL: AI systems already out-persuade expert humans (from the UK AISI evals).
- INCIDENT: The labs' cybersecurity is crap: in July, white hats used Claude to break into OpenAI, reach an employee's Codex account and open a pull request in its internal monorepo. "Win the race" now means winning for 1 week before all your algorithmic secrets leak out.
A substitute for evals is to simply monitor and report whatever bad stuff (and impressive stuff) people are trying to do with deployed systems. For example, we know from Anthropic's threat report that some scientists, including one affiliated with a military institute, used Claude for gain-of-function research that could support bioweapons. AI compa