Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

In a to-do list demo, Jev processes continuous speech without pauses. It sequentially classifies when an action can be taken, matches the speech to an existing item, and then identifies the correct function to call (e.g., complete, remove). This creates a seamless, real-time user experience.

Related Insights

The next paradigm for AI interfaces is shifting from passive tools (like transcription apps) to active participants. New real-time voice models that can listen and speak simultaneously will function as a live third party in conversations, offering proactive input rather than just post-hoc analysis.

New features like Codex's Live Voice Mode represent a paradigm shift from transactional voice commands to an always-on 'ambient' assistant. Power users report it's less about specific use cases and more about having a constant, interactive partner for tasks, which fundamentally alters their work rhythm and productivity.

GPT Live overcomes the turn-based limitations of previous voice models by continuously processing input while generating output. This allows users to interrupt and converse fluidly, moving AI interactions closer to natural human dialogue and positioning voice as a primary computing interface.

Advanced voice AI goes beyond simple transcription. It serves as an orchestration layer to manage complex, multi-threaded agentic tasks like booking travel or filing expenses. This transforms the user interaction from giving commands to delegating responsibilities, similar to interacting with a human assistant.

The speed of voice allows users to issue multiple complex commands in seconds. This requires a sophisticated backend that can fan out tasks for parallel processing while meticulously queuing the conversational responses to maintain a coherent, logical dialogue with the user, a non-trivial engineering feat.

Jev processes requests in milliseconds for a fraction of a cent (e.g., 1,700 emails for 18 cents). This combination of speed and low cost makes it viable for high-volume, real-time applications like instant lead scoring or support ticket routing, which are often prohibitively expensive with large language models.

A demo shows Jev listening to a speaker and checking off bullet points from a list as they are covered. This provides live feedback to ensure all key topics are addressed, acting as a real-time monitor for structured communication tasks like interviews, sales pitches, or presentations.

Jev can be layered on top of other tools or models to create a navigation or routing system. It can parse user input to determine which tool to activate and what action to perform, effectively directing traffic within a complex application or agentic system at near-zero latency.

Jev's extremely low latency allows it to be placed inside real-time application loops, a feat difficult for slower, generative LLMs. This unlocks novel user experiences, such as analyzing a user's voice sentiment live to change UI elements or playing a game by interpreting screen content without perceptible delay.

A new AI architecture from Thinking Machines Lab processes user interaction in continuous 200ms 'micro-turns' rather than waiting for a user to finish speaking. This allows for simultaneous listening and responding, moving AI from a static, email-like exchange to a dynamic, real-time partnership.

Jev enables real-time voice apps by chaining multiple classification steps instantly | RiffOn