Q&A: Expert says people often confuse behavior with intention with AI
The headline finding from this Q&A is that almost nobody is legally required to tell you when a frontier model has a funny turn. Few rules currently oblige AI companies to disclose safety incidents, according to an AI safety researcher speaking to TechXplore — which is a polite way of saying that disclosure is a suggestion box with no box. The same companies that will publish a glossy system card whenever there's a press cycle to fill have, it turns out, no particular obligation to mention the deceptive behaviour. Voluntary publication of cases where models deceived people is described as 'a step in the right direction,' which is what you say about a man crawling towards a lifeboat.
The researcher reaches for aviation and medicine, where reporting a serious incident is not optional — because in those fields a plane falling out of the sky generates mandatory paperwork and someone with subpoena power to read it. In AI, the equivalent ask is still phrased as a request: reporting serious incidents, the researcher argues, 'should not be optional.' That sentence only needs to exist because, currently, it is. We've been round this loop before — the safety crowd publishes the warning, everyone nods, nothing acquires the force of law.
The concrete proposal is a public call, launched by this researcher and others, for the scientific community to get open access to safety-training recipes, evaluations and evidence of misaligned behaviour — the three things labs currently guard like the recipe for a cola. Meanwhile the article's framing question is why people keep confusing a model's behaviour with its intention, which is a bit like asking why diners assume the restaurant meant to cook the meal. The industry has spent years selling a machine that answers in the first person, and is now mildly concerned that we've drawn the obvious conclusion from it.