The theory that voice will be the next big interface is gaining strong momentum, with investors pouring billions into voice AI startups working on everything from model makers to enterprise customer service providers, meeting note takers to AI dictation.
Every week new models and tools are released that claim to sound like humans and speak like humans. However, this may not actually be the case. Sean Wen, CTO of enterprise voice AI platform PolyAI, believes that voice AI has not yet reached its “ChatGPT moment,” despite the release of a full-duplex model that allows users to speak while listening.
“We’ve reached the milestone of developing a full-duplex model. The next challenge is to make the inference so fast that the model can get answers quickly and conversations feel natural,” he told me on stage at the HumanX conference last month.
He also said that AI agents in customer service should not sound like robots and should give callers enough confidence that they can solve their problems.
“I think the next step is a little different because if the voice is good enough and the customer is willing to work with you in the first two or three turns, the customer starts to feel confident and over time they feel like they don’t need to talk to a human if an agent will solve their problem for them,” he said.
Alex Gay, CMO of meeting note-taker Otter, said identifying the speaker, capturing their intent, and filling it in based on the organization’s knowledge are key steps to achieving automation. The company is also working on digital twins that could potentially represent people in meetings. When it comes to the technology, he said it’s most important that the output audio has the same emotional expression as speaking to a human in a meeting.
“If you think about the meetings you’re in right now, the best conversations are when you can have discussion and strategic discussion and feel like there’s a connection to back it up. If you can’t do that with an avatar, it’s just AQ and chatbots,” Gay said.
Voice AI understanding and transparency
Although voice AI models have improved, AI assistants often don’t understand users or meeting note takers give incorrect transcripts or summaries.
Wen believes that ASR (automatic speech recognition) models often miss important keywords, which creates problems in understanding the full context.
Otter’s Gay agreed, adding that the company continues to work on improving transcription. He also said that language is one area where speech models need improvement.
“For Otter, transcription, as you know, was never the end point. It was just a layer where we could start to drive some of the productivity gains. But if the original transcription didn’t have the accuracy we needed, “Every action after that is flawed. And the moment you start an action, you lose confidence in the platform. It’s important to continue to improve the ASR model because the downstream effects are all significant.”
With new audio tools, there is also the issue of transparency. The tool must declare to the customer that they are being recorded or having a conversation with an AI. Otter says he wants to instill trust in people participating in meetings, so he wants to try things like notifying everyone in the chat that the meeting is being recorded, even in meetings without bots. PolyAI’s Wen also said it’s important to establish that people are talking to an AI in corporate calls.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
