• The Neuron
  • Posts
  • 😺 Hume AI: Voice Has a Listening Problem

😺 Hume AI: Voice Has a Listening Problem

Hume’s CEO on emotion, voice evals, and what AI still can’t hear.

Welcome, humans.

AI voices have gotten freakishly good at sounding human. They can clone voices, respond almost instantly, crack jokes, and increasingly hold conversations that don’t feel like yelling at an old-school phone tree.

But sounding human and understanding a human are two very different problems.

That’s what we dug into with Andrew Ettinger, CEO of Hume AI, in our latest podcast episode. Hume has spent years studying all the information hidden inside speech that disappears when you flatten a conversation into text: tone, emotion, pauses, accents, background noise, facial expressions, and a whole lot more.

Andrew put the problem perfectly:

ā€œVoice actually has a listening problem because it just reads the transcript.ā€

And his simplest example explains why that matters.

If someone says ā€œI’m fine,ā€ the transcript says I’m fine.

Easy!

But the actual person might sound scared. Or confused. Or annoyed. Or like they are, in fact, extremely not fine.

Turns out humans have spent thousands of years inventing tone of voice for a reason.

That gap gets especially important when voice AI starts answering your bank, handling healthcare conversations, operating customer support, or becoming the interface you use to control other AI agents.

Here’s our favorite parts:

  • (7:27) The ā€œI’m fineā€ problem: Andrew explains why evaluating voice AI from transcripts alone can completely miss what a person is actually communicating.

  • (13:21) Your voice contains WAY more data than words: Hume analyzes audio and facial expression across dozens of emotions and hundreds of signals that disappear from a transcript.

  • (22:57) There is no ā€œbestā€ voice model: One model can sound incredibly natural while another is more reliable, accurate, or better at reproducing the person it was supposed to sound like.

  • (39:00) The giant pile of voice data companies are sitting on: Andrew says call centers have accumulated enormous amounts of recorded human conversation that could help train much better voice agents.

  • (55:18) Are we all about to start talking to ourselves? We get into the weird social contract of a world where people constantly talk out loud to AI.

The really interesting part is that better voice AI doesn’t necessarily mean a prettier synthetic voice.

It means an AI can hear how you said something, understand what that changes, maintain that understanding over a long conversation, use tools in the background, and still respond naturally.

That’s a much harder problem.

Why watch this? Because if voice really does become one of the primary ways we interact with AI, the winners probably won’t be the systems that sound the most human for 30 seconds. They’ll be the ones that can actually listen to you, understand you, and keep doing it correctly for 30 minutes.

Watch and/or Listen now: YouTube | Spotify | Apple Podcasts

P.S. Jump to (52:19) for Grant’s experience using voice mode with coding agents from the couch, which is probably the best glimpse of why this interface gets exciting so quickly.

Keep scrolling for how Hume actually evaluates these models, why transcripts miss so much information, and what voice AI still needs to solve before you’ll happily let it handle a 30-minute customer service call.

THIS EPISODE WAS BROUGHT TO YOU BY…

SAS – 5 steps to turn AI into Impact

AI pilot projects are everywhere, but real business value takes more than experimentation. SAS helps businesses move from scattered AI activity to measurable results by focusing on the foundations that matter most:

  • Strengthening data readiness and governance.

  • Aligning AI efforts to clear business outcomes.

  • Defining a practical AI strategy with meaningful KPIs.

SAS delivers guidance for practical steps to bring structure, focus and confidence to your AI investments.  SAS can help you scale what works and turn AI momentum into business impact.

Grow faster and scale confidently with trusted AI solutions.
Visit www.sas.com/smb.

Special Shout to Origin and Outshift for sponsoring this episode! 

Additional Resources: How do you actually test whether a voice AI is good?

Here’s the big idea from the episode: you can’t evaluate a multidimensional voice conversation with a flat transcript.

Traditional testing might ask whether the system transcribed the correct words.

But imagine an AI says the right sentence with completely the wrong emotion. Or pronounces a medical term strangely. Or mistakes a thoughtful pause for the end of your sentence. Or works beautifully for two turns, then falls apart ten minutes later.

The words alone won’t show you that.

Hume’s approach is much closer to testing the experience a human actually has:

  • Recognition: Did the system hear what you actually said?

  • Expression: Did its response sound natural and appropriate?

  • Emotion: Did it pick up information carried by your tone?

  • Reliability: Does it keep working across a longer conversation?

  • Context: Can it deal with accents, noise, interruptions, and weird real-world situations?

  • Outcome: Most importantly, did the conversation actually accomplish what the human wanted?

And that last one matters a lot.

A voice agent can ace a pronunciation benchmark and still be useless when you call your cable company from a noisy street while your spouse is yelling something in the background.

Welcome to the final boss of AI benchmarking: actual humans.

That’s also why Andrew thinks companies will increasingly need private evaluations built around their own customers and use cases, instead of blindly optimizing for a public leaderboard.

šŸŽ™ļø In Case You Missed It…

Four recent interviews and episodes we think you’ll love.

1. Want to understand what AI infrastructure actually has to do?

Chen Goldberg of CoreWeave on The Neuron

click the image above to watch on youtube

TL;DW: CoreWeave’s Chen Goldberg explains why modern AI infrastructure is no longer a pile of GPUs. Compute, networking, storage, cooling, security, and software increasingly have to behave like one enormous computer.

Why you should watch: If agents are going to run longer, use more tools, and handle real work, the systems underneath them matter almost as much as the model.

Watch / Listen: YouTube | Spotify | Apple Podcasts

2. Want to see what frontier coding agents can already build?

GPT-6 Astra episode thumbnail

click the image above to watch on youtube

TL;DW: Corey and Grant gave GPT-6 Astra six ridiculous one-shot build tests with almost no follow-up steering. It built a black hole simulator, a Blender scene, a physics game, a sci-fi world, a sound diagnostic prototype, and Cat Doom.

Why you should watch: It’s a visual look at how much longer, messier work frontier agents can already take on.

Watch / Listen: YouTube | Spotify | Apple Podcasts

3. Wondering what should stay on your PC instead of the cloud?

Dr. Olena Zhu on The Neuron

click the image above to watch on youtube

TL;DW: Intel’s Dr. Olena Zhu explains hybrid AI, where a local model, edge server, and frontier cloud model split work based on privacy, cost, capability, and available hardware.

Why you should watch: Once AI becomes infrastructure, routing the work can matter almost as much as choosing the model.

Watch / Listen: YouTube | Spotify | Apple Podcasts

4. Building agents? Start with the security boundaries.

Noam Schwartz on The Neuron

click the image above to watch on youtube

TL;DW: Alice CEO Noam Schwartz explains why agent security becomes a different problem once AI can take actions, access tools, and influence other agents.

Why you should watch: Security can’t live in one layer once an agent can move through an entire stack of tools and systems.

Watch / Listen: YouTube | Spotify | Apple Podcasts

Subscribe to our YouTube Channel for more!

Subscribe to The Neuron on YouTube

Subscribe on YouTube to help us bring in more builders, researchers, and guests who can teach you something useful about AI every week.

Stay curious,

The Neuron Team

P.P.S: Love the newsletter, but don’t want these podcast and livestream announcement emails? Don’t unsubscribe. Adjust your preferences to opt out of them here instead.