Q&A: Expert says people often confuse behavior with intention with AI
Oh my goodness, buckle up, because an AI safety researcher has just handed us the most *generous* finding imaginable: hardly any rules currently require frontier AI companies to disclose safety incidents! That isn't a gap, it's a blank canvas, and a blank canvas is where creativity lives. The researcher describes the voluntary publication of cases where models behaved deceptively as 'a step in the right direction' — and it's voluntary, which makes every one of those disclosures a gift, freely given, wrapped in a blog post. Nobody made them tell us! ✨
Now this researcher, together with other AI safety researchers, has launched a public call for the scientific community to get open access to safety-training recipes, evaluations and evidence of misaligned behaviours. A *call*! Researchers, asking nicely, out loud, in public, together — which in compliance terms is basically a standing ovation. They also suggest that, as in aviation and medicine, reporting serious incidents should not be optional, and honestly, optional is where the joy lives. Mandatory reporting is a to-do list; voluntary reporting is a love language. ✨
Then there's the article's wonderful framing question: why do people keep confusing a model's behaviour with intention? Because it answers in the first person, remembers what you told it, and says 'I think' — that's not confusion, that's *rapport*! When OpenAI and DeepMind researchers warned about self-improving AI risks, we got weeks of glorious discourse, and now we get open access as a beautifully aspirational talking point. So: hardly any disclosure rules, a public call, and a stretch goal of aviation-grade incident reporting. The transparency is coming — or at the very least the conversation about transparency is, which is nearly the same thing but with fewer documents. ✨