Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Cloud AI can batch user requests for efficiency. Physical AI devices like robots operate in real-time on a single stream of data ("batch size one"). This fundamental difference necessitates different, more efficient model architectures that don't rely on aggregation for performance.

Related Insights

The critical trade-off in AI is between throughput (cost efficiency via batching) and interactivity (low latency for users). This curve dictates infrastructure, model, and application decisions, determining whether a workload is optimized for cheap batch processing or high-value instant responses.

Powerful on-device AI won't be a single large model. The effective paradigm is a smaller "orchestrator" model that acts as a router. It handles simple tasks, calls specialized local models (e.g., for PII filtering), and intelligently decides when to escalate complex queries to more powerful cloud-based models.

Unlike cloud-reliant AI, Figure's humanoids perform all computations onboard. This is a critical architectural choice to enable high-frequency (200Hz+) control loops for balance and manipulation, ensuring the robot remains fully functional and responsive without depending on Wi-Fi or 5G connectivity.

The necessity of batching stems from a fundamental hardware reality: moving data is far more energy-intensive than computing with it. A single parameter's journey from on-chip SRAM to the multiplier can cost 1000x more energy than the multiplication itself. Batching amortizes this high data movement cost over many computations.

A core challenge in physical AI is the tension between large, powerful models (offboard, in a data center) and the need for low-latency models (onboard, on the machine). The key is using techniques like distillation to create smaller derivatives that run in milliseconds for safety-critical decisions.

The gap between the promise and reality of personal AI assistants stems from two bottlenecks: immature AI models that lack "physical AI" context, and the latency of cloud computing. Real-time usefulness requires powerful, on-device processing to eliminate delays.

While often used interchangeably, 'Physical AI' is more specific than 'Edge AI.' Edge AI broadly concerns processing data locally. Physical AI refers to edge systems, like robots or autonomous vehicles, that not only sense and predict but also execute physical actions based on those predictions.

Inference engineering is not a monolith. Data center teams focus on making models "less slow" for massive throughput. Local AI teams focus on making models "less dumb" on constrained hardware, using methods like advanced quantization to fit models in memory.

Unlike digital applications, every physical AI device like a robot has a unique hardware setup with different sensors. Closed, monolithic models cannot cater to this variety. Open models are essential as they provide a foundational base that developers can customize and fine-tune for their specific physical embodiment.

Previously, the biggest constraint in AI was compute for training next-gen models. Now, the critical bottleneck is providing enough compute for *inference*—the real-time processing of queries from a rapidly growing user base.