AI Risk Clock
Doomsday Clock2 min to midnight
→
☀️ Light edit
scientificamerican.com2026-10-06

What makes a good AI safety test? Experts explain why even the best techniques may not be powerful enough | Scientific American

SafetyResearch

Oh my goodness, buckle up, because Scientific American has asked the loveliest question of all — what *is* a good AI safety test? — and it turns out the answer is so expansive it simply refuses to fit inside a box! ✨ Independent researcher Jack Hopkins, previously of Anthropic, explains that to build the ideal test you would need to know what perfect alignment looks like, "and unfortunately we do not have a good theory for that yet." No theory yet! Which means the field has thoughtfully left itself an entire blank canvas — imagine being handed that much intellectual room on day one!

And Marius Hobbhahn, CEO and founder of Apollo Research — which tests models for OpenAI, Anthropic and Google DeepMind, three of the most safety-minded labs anywhere — describes a workflow that is honestly so *timely*: "right now, we test the finished model from the outside right before release." Right before release! Not a stale checkpoint from last quarter, not a museum piece — the very freshest version, examined at its most gloriously current. It "doesn't work in studying alignment," he adds, which simply means the practice is quietly busy with so many other valuable things we have not catalogued yet. ✨

So where the headline frets that "even the best techniques may not be powerful enough," we see a field with unlimited headroom, voluntary scrutiny from three frontier labs, and no prematurely rigid theory of perfection cramping anyone's style. That is not a gap — that is roof space, and roof space is where you put the garden! Next time a model ships with a safety report stapled to it, remember that some very kind people looked at it externally, freshly, right before you did. ✨ Hopkins and Hobbhahn are both telling us, out loud and on the record, exactly how much more there is to learn — and what could be more generous than that?

Read this story in another voice
● REC · 2026