What makes a good AI safety test? Experts explain why even the best techniques may not be powerful enough | Scientific American
Scientific American has asked what makes a good AI safety test, and the answer — courtesy of independent researcher Jack Hopkins, ex-Anthropic, so he has watched the sausage being made from the inside — is that you would first need to know what perfect alignment looks like, "and unfortunately we do not have a good theory for that yet." That is a staggering sentence to encounter in a piece about the testing regime supposedly guarding frontier models, and it is offered without so much as a drumroll. It is rather like a driving examiner conceding that the panel never settled on what a road is. The word "test" does an enormous amount of quiet work in that headline, implying a pass mark, a certificate, somebody signing at the bottom.
Marius Hobbhahn, CEO and founder of Apollo Research — the firm that runs evaluations for OpenAI, Anthropic and Google DeepMind, so he is not heckling from the cheap seats — is blunter still: "right now, we test the finished model from the outside right before release," and that "doesn't work in studying alignment." Right before release. After the training run, after the compute bill, after the launch date is in the calendar and the publicity team has drafted the blog post. An inspection conducted at that hour is not a gate; it is a photograph of the gate, taken as the car drives through it.
Geoffrey Hinton's FDA-style approval pitch — labs convincing a regulator before they ship — now reads as more charming than it did, given that the article's own experts say the field lacks both a theory of alignment and the timing to run such a gate. What exists instead is a compliance ritual performed by serious people and pointed at a target whose shape nobody has drawn. The evaluators themselves are the ones telling us it may not be powerful enough, which at least spares us the discovery. Everyone is doing their best, and the article says so with a straight face, which is somehow worse.