Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

With internet data fully exploited, the next frontier for training large-scale AI models is scientific experimentation. The scientific method, using nature as a verifier, can generate a virtually endless stream of novel, high-value data ('tokens') to advance AI reasoning capabilities.

Related Insights

Unlike language models trained on the internet, AI for materials science overcomes data scarcity and unreliability (e.g., conflicting literature) with a closed loop. The system actively directs experiments, analyzes grounded results for patterns, and uses that new data to drive the next cycle.

Future progress in biology requires moving beyond static models. The new paradigm involves an AI that reasons over hypotheses, prioritizes experiments, learns from the empirical outcomes, and updates its internal world model. This creates a scalable, closed-loop system for scientific discovery.

The next leap in biotech moves beyond applying AI to existing data. CZI pioneers a model where 'frontier biology' and 'frontier AI' are developed in tandem. Experiments are now designed specifically to generate novel data that will ground and improve future AI models, creating a virtuous feedback loop.

Early AI models advanced by scraping web text and code. The next revolution, especially in "AI for science," requires overcoming a major hurdle: consolidating and formatting the world's vast but fragmented scientific data across disciplines like chemistry and materials science for model training.

The internet is an insufficient training ground for scientific AI because most crucial information—including failed experiments, negative data, and nuanced procedural details—is never published. This undocumented knowledge, what scientists call "good hands," represents a major data bottleneck for building truly intelligent scientific models.

Instead of generating data for human analysis, Mark Zuckerberg advocates a new approach: scientists should prioritize creating novel tools and experiments specifically to generate data that will train and improve AI models. The goal shifts from direct human insight to creating smarter AI that makes novel discoveries.

Static data scraped from the web is becoming less central to AI training. The new frontier is "dynamic data," where models learn through trial-and-error in synthetic environments (like solving math problems), effectively creating their own training material via reinforcement learning.

The ultimate goal isn't just modeling specific systems (like protein folding), but automating the entire scientific method. This involves AI generating hypotheses, choosing experiments, analyzing results, and updating a 'world model' of a domain, creating a continuous loop of discovery.

Current LLMs fail at science because they lack the ability to iterate. True scientific inquiry is a loop: form a hypothesis, conduct an experiment, analyze the result (even if incorrect), and refine. AI needs this same iterative capability with the real world to make genuine discoveries.

The founder of AI and robotics firm Medra argues that scientific progress is not limited by a lack of ideas or AI-generated hypotheses. Instead, the critical constraint is the physical capacity to test these ideas and generate high-quality data to train better AI models.

Scientific Experiments Are the 'Infinite Token Generator' for Post-Internet AI | RiffOn