Brief

Voice AI leaders say the technology still lacks a ChatGPT‑style breakthrough

Industry executives argue that voice AI must improve speed, reasoning and transcription accuracy before achieving mainstream adoption.

By Felo News Desk · Published

At the HumanX conference last month, PolyAI chief technology officer Shawn Wen said voice AI has not yet reached a "ChatGPT moment," despite recent advances such as full‑duplex models that can speak while listening. Wen told TechCrunch that the next hurdle is making reasoning fast enough for natural‑sounding conversations.

What happened

Wen highlighted the development of full‑duplex voice models as a milestone, but emphasized that rapid reasoning is required for callers to feel confident that an AI agent can solve problems without human intervention. He suggested that once voice quality is sufficient for the first two or three turns, users may begin to rely on the agent instead of a human.

What the reports add

Alex Gay, chief marketing officer of Otter, added that speaker identification, intent capture and integration with organizational knowledge are critical for automating meeting workflows. Gay said Otter is also exploring digital‑twin avatars that can convey the same emotive expressions as a human speaker, warning that a lack of relational nuance would reduce interactions to a simple Q&A chatbot.

What was said

"We have reached the milestone of developing full‑duplex models. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural," Wen told the conference, as reported by TechCrunch.

Gay explained, "If you think about the meetings that you’re in right now, the best conversations that you have are where you can have debate, and strategic discussions, and when you feel like there’s a relationship that underpins it. If you aren’t able to have that with an avatar, then it’s just a q and a chatbot." He also warned that inaccurate automatic speech recognition (ASR) can erode trust: "If your original transcription didn’t have the accuracy that you needed, all of the follow‑up actions become flawed… It is critical for us to continue to improve that ASR model because all of the downstream impacts are significant."

How it came about

Investors have poured billions into voice AI startups covering model creation, enterprise customer‑service platforms, meeting note‑taking and AI‑powered dictation. Weekly releases claim human‑like speech and conversation, yet Wen and Gay argue that current models still fall short on speed, reasoning and transcription fidelity, limiting broader adoption.

Key facts

  • PolyAI CTO Shawn Wen says voice AI lacks a ChatGPT‑style breakthrough because reasoning speed is still insufficient. (techcrunch.com)
  • Otter CMO Alex Gay stresses the need for accurate speaker identification, intent capture and improved ASR to build trust in voice assistants. (techcrunch.com)
  • Full‑duplex voice models, which can speak while listening, have been developed but are not yet considered a turning point. (techcrunch.com)

Sources

  • [1] techcrunch.com — originally reported as “These execs think voice AI hasn't reached its ChatGPT moment yet”

Earlier coverage

More from Technology

Felo News, House 42, Bridge Colony, Kot Lakhpat, Lahore, Pakistan
+92 308 4354717 · felopronews@gmail.com