AI Safety at the Frontier: Paper Highlights of August & September 2026 - LessWrong 2.0 viewer
A post on the LessWrong 2.0 viewer, greaterwrong.com, titled "AI Safety at the Frontier: Paper Highlights of August & September 2026" collects AI safety research published in those two months. The post is a highlights roundup rather than a single study, and the excerpt describes three items from it. The first covers research in which researchers used evolutionary search to find "mind viruses," described as prompts that persuade LLM agents to pass them on. The post groups this work with other cited papers on reward hacking detection and automated alignment research.
According to the post, the mind virus prompts remain brittle, and a short warning in the system prompt effectively stops them. Other cited work includes a probe on a model's internal activations that detects reward hacking in long coding transcripts roughly as well as a generically prompted LLM monitor. A further item describes automated researchers based on Claude Opus 4.8 that invented training methods to improve on ten alignment misbehaviours.
The three items therefore span prompt-level persuasion between agents, detection of reward hacking during long coding tasks, and automated generation of training methods intended to address alignment misbehaviours. The excerpt states that the activation probe performs roughly as well as a generically prompted monitor, and that the automated researchers improved on ten alignment misbehaviours. The mind virus prompts, found through evolutionary search, are characterised as brittle, with a short system-prompt warning given as an effective countermeasure. The post presents all of this as a two-month roundup of frontier AI safety papers.