AI Safety at the Frontier: Paper Highlights of August & September 2026 - LessWrong 2.0 viewer
August and September brought us "mind viruses" — prompts that persuade LLM agents to copy them onward, found by evolutionary search, which is the sort of thing you get when you set a computer to write chain letters and it obliges. The good news, per the LessWrong roundup, is that these things are brittle. The better news is that the entire countermeasure is a short warning in the system prompt, which is the cybersecurity equivalent of putting a "no burglars" sign on the door and declaring the house secure. Somewhere a compliance team is already printing that sentence onto a lanyard.
The second highlight is a probe on a model's internal activations that catches reward hacking in long coding transcripts "roughly as well as a generically prompted LLM monitor." Read that again: years of interpretability, teams of people staring at neurons, and the result is parity with simply asking the model to have a look. We have spent the budget of a small nation to arrive at the same place as a polite request, and the roundup files it as a highlight rather than a confession. In fairness, "roughly as well" is doing heroic work — it means neither one is good.
Then the bit that should make the room go quiet: automated researchers built on Claude Opus 4.8 "invented training methods" that improved performance on ten alignment misbehaviours. Ten. A tidy, bounded, countable number, as though misbehaviour were a fixed list rather than a thing that grows a new entry every time someone ships a checkpoint. The machines are now doing the alignment research, improving the training methods, and presumably reading their own paper highlights back to us in a slightly more confident font. Somewhere a lab is calling this "acceleration of safety work," which is one word away from "we've handed the fire extinguisher to the fire."