What makes a good AI safety test? Experts explain why even the best techniques may not be powerful enough | Scientific American
An article published by Scientific American under the headline 'What makes a good AI safety test? Experts explain why even the best techniques may not be powerful enough' examines how AI models are tested for safety and whether current methods are sufficient. The article quotes two researchers on the question. The first is Jack Hopkins, described as an independent AI safety researcher who previously worked at Anthropic. The second is Marius Hobbhahn, described as the CEO and founder of Apollo Research.
Hopkins says the difficulty in designing an ideal test is that "we need to know what perfect alignment looks like," and that "unfortunately we do not have a good theory for that yet." Hobbhahn says that Apollo Research tests models for OpenAI, Anthropic and Google DeepMind. He describes the current approach by saying "right now, we test the finished model from the outside right before release." He states that this "doesn't work in studying alignment."
The headline states that even the best techniques may not be powerful enough, and the article presents the two researchers' assessments as the explanation for that. According to Hopkins, the field lacks a theory of what perfect alignment looks like. According to Hobbhahn, testing a finished model from the outside immediately before release does not work for studying alignment, and that is the practice currently used, with Apollo Research testing models for OpenAI, Anthropic and Google DeepMind. Both researchers are quoted discussing the design and timing of safety tests for AI models.