Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A digital twin trained only on a single lab's limited data is not extendable or robust. The future is building on "foundation models" trained on massive, public consortium data, which allows you to transfer learning and strengthen specific models.

Related Insights

To build generalist robots, the most effective approach is pre-training foundation models on internet-scale video datasets, not just simulation or tele-operated data. This vast, diverse data provides a deep, implicit understanding of physics and object interaction that is impossible to replicate in controlled environments, enabling true generalization.

The Physical Intelligence thesis is that a foundation model learning from diverse data can achieve a "physical understanding" of the world, making it easier to adapt to new tasks than building single-purpose robots from scratch. Generality leverages broader data, which is ultimately a more scalable approach.

The primary bottleneck for creating powerful foundation models in biology is the lack of clean, large-scale experimental data—orders of magnitude less than what's available for LLMs. This creates a major opportunity for "data foundries" that use robotic labs to generate high-quality biological data at scale.

Instead of relying on opaque model weights, continual learning is more reliably achieved by having AI build explicit, external 'world models' like knowledge graphs. This approach makes the model's understanding inspectable and correctable by humans, enabling more robust causal analysis.

Adopting the true 'foundation model ethos' from LLMs is difficult for roboticists. It means a warehouse automation company should collect data from kitchen robots. This breadth, while seemingly unrelated, builds a generalist model that better handles the weird edge cases in the target domain than a narrowly trained specialist model.

The core weakness of Transformers is their static nature. They are trained in a lab on a snapshot of data and then deployed. They cannot adapt to new events, tools, or user tasks without a full retraining cycle, making true continuous learning at test time impossible with the current architecture.

The true cost of fine-tuning isn't the initial training but the ongoing maintenance. Base foundation models experience significant capability improvements every 2-3 months. This pace means a custom fine-tuned model can quickly fall behind, forcing a continuous and expensive re-tuning cycle.

Building one centralized AI model is a legacy approach that creates a massive single point of failure. The future requires a multi-layered, agentic system where specialized models are continuously orchestrated, providing checks and balances for a more resilient, antifragile ecosystem.

Instead of costly proprietary data generation, Turbine focused on the 'unsexy' work of combining many different public and partner datasets. This capital-efficient approach forced them to build an AI model architected for generalization and data efficiency from the very beginning.

Unlearn.ai found that scaling digital twins from CNS to oncology isn't about parameter changes. Radically different data structures—like oncology's hierarchy of rare diseases and complex treatment histories—demand entirely new modeling approaches, unlike the more siloed data found in CNS trials.