OpenAI bets voice will become AI's primary interface with new models - Axios
Oh, brilliant. OpenAI, in its infinite wisdom and despite its revolving door of executives, has decided that the future of human-computer interaction is voice. The company is rolling out new voice models for ChatGPT that route queries to its latest frontier text models, with the explicit goal of making voice the primary interface for AI. Because why let people type their requests when they could whisper sweet nothings to a model that might hallucinate their medical advice?
The logic is that voice is more natural than text, as if naturalness is the priority over accuracy. These frontier models are already notorious for fabricating sources and facts with panache, and now they'll be talking back to you in a calm, reassuring tone. Imagine Amazon Alexa but with the underlying intelligence of a GPT-5 that sometimes thinks you're a toaster. OpenAI's bet is that users will prefer a spoken conversation over reading, which certainly explains why they keep pushing voice despite the documented problems with audible deception. The new models ensure that your voice query is processed by the absolute best text models, so you'll be getting premium nonsense delivered right to your ears.
But here's the darkest part: voice models are harder to monitor. There will be no way to quickly scroll and verify a spoken claim. You'll have to ask it to repeat itself, and by then it might have changed its story. OpenAI is essentially training us to accept verbal output from systems that cannot be held accountable. Meanwhile, regulators are still wondering whether to send a strongly worded letter or a tweet. So go ahead, set voice as your primary interface. What could possibly go wrong? Your new AI friend is waiting. It just might not be telling the truth. And it'll sound so confident doing it.