Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

LiveKit invested heavily in fixing "backchanneling," where an AI mistakes human interjections like "uh-huh" for an interruption. By solving these nuanced, infuriatingly non-human flaws, they provide immense value to developers trying to build natural-feeling agents, building more loyalty than a major feature might.

Related Insights

The primary reason voice assistants feel robotic is their failure to process audio while speaking. They get confused by simple interjections like "yeah" or attempts to interrupt. OpenAI's new "BIDI" model aims to solve this by listening and updating its response in real-time for a more natural conversation.

While Genspark's calling agent can successfully complete a task and provide a transcript, its noticeable audio delays and awkward handling of interruptions highlight a key weakness. Current voice AI struggles with the subtle, real-time cadence of human conversation, which remains a barrier to broader adoption.

Meet Sona's team initially implemented a sophisticated, real-time conversational AI that could interrupt users to feel more natural. They discovered through user feedback that this was overwhelming and stressful. They deliberately simplified the experience, adding user-controlled pauses and prioritizing user comfort over technical wizardry.

Building loyalty with AI isn't about the technology, but the trust it engenders. Consumers, especially younger generations, will abandon AI after one bad experience. Providing a transparent and easy option to connect with a human is critical for adoption and preventing long-term brand damage.

For high-stakes, long-duration calls (e.g., remote patient monitoring), AI cannot be a rigid phone tree. To gain the trust of users like elderly patients, the AI must be able to navigate tangential personal stories—'hear about their grandchild'—before it can effectively guide them through a complex task. This human-centric approach is non-negotiable.

Even after disclosing that an agent is an AI, prioritizing a human-like conversational experience is critical. Users quickly forget they're talking to a machine if the interaction is natural, which reduces friction and makes the automation more effective and accepted.

Contrary to fears of customer backlash, data from Bret Taylor's company Sierra shows that AI agents identifying themselves as AI—and even admitting they can make mistakes—builds trust. This transparency, combined with AI's patience and consistency, often results in customer satisfaction scores that are higher than those for previous human interactions.

The team's breakthrough moment wasn't perfect voice replication, but when their AI model first laughed. They realized that human-like imperfections—laughter, pauses, "ums"—were the critical elements that made the user experience feel genuinely human and believable, leading to their first viral moment on Hacker News.

By meticulously prompting the AI to use an informal, lowercase, and sometimes profane tone, Lindy makes its mistakes feel more human and less jarring. When the AI says 'oh, shit. You're right,' it 'takes the edge off the fuck up,' building user trust and rapport.

For personal AI agents like OpenClaw, the conversational interface—feeling like you're texting a person—accounts for the vast majority of user adoption and value. This emotional, personal connection is far more important than the agent's technical capabilities, like self-modification or its skills directory.