Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A new AI model named Griffin by Tavis can render a lifelike person that reacts in real-time on a video call. In a company test, nearly half of participants believed they were interacting with a human. This low-latency, responsive interaction marks a significant milestone in generative AI, effectively passing a practical visual Turing test.

Related Insights

The next paradigm for AI interfaces is shifting from passive tools (like transcription apps) to active participants. New real-time voice models that can listen and speak simultaneously will function as a live third party in conversations, offering proactive input rather than just post-hoc analysis.

Tavis's AI model was believed to be human by 48% of testers, but 15% also thought a real human was AI. This highlights that a true measure of passing the Turing Test isn't hitting an absolute 50% but significantly outperforming the human 'false positive' rate, accounting for inherent user skepticism.

While the AI avatar achieved a strong physical likeness, especially in profile, it failed to render nuanced emotions convincingly. The host described a scene of her laughing as "100% uncanny valley," indicating that current models still struggle to cross the emotional authenticity barrier needed for believable human characters.

Muse uses on-demand image generation to create and animate its avatar, which even gets a tiny laptop when working. This delightful, dynamic interaction transforms the agent from a tool into a character, fostering a stronger emotional connection and user engagement than a static UI.

The long-held standard for machine intelligence, the Turing Test, is now routinely passed by commercial AI models. Its failure as a good measure of general intelligence has rendered it obsolete, demonstrating that facility with language does not equate to the broader cognitive capabilities once assumed.

The next wave of AI assistants focuses on "interaction" or "bi-directional" models that can process information and respond in real-time, allowing users to interrupt them naturally. Startups like Thinking Machines Lab are competing directly with giants like OpenAI to create a more fluid, human-like conversational experience, moving beyond today's turn-based models.

The next frontier for visual intelligence is twofold: creating truly multimodal models that retain long-term context of user interactions without re-prompting, and developing real-time generation. Real-time capabilities are crucial for creating duplex interactions and enabling robots to perceive and act instantly.

A viral demo of Kling AI's "motion transfer" feature shows a user's live movements being perfectly mirrored by a photorealistic avatar in real-time. This capability goes beyond static deepfakes, introducing live, user-controlled synthetic video that drastically blurs the line between reality and AI generation.

The next frontier for conversational AI is not just better text, but "Generative UI"—the ability to respond with interactive components. Instead of describing the weather, an AI can present a weather widget, merging the flexibility of chat with the richness of a graphical interface.

AI video is evolving from passive generation to active engagement. Synthesia's new products focus on the intersection of video and AI agents, allowing users to, for example, watch a training video and then enter a role-playing simulation with an AI to test their comprehension.

Tavis's 'Griffin' Avatar Suggests the Visual Turing Test Has Been Passed for Live Video Calls | RiffOn