Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Joon Sung Park notes they are observing the beginnings of a scaling law for simulation. As they ingest more compute and high-quality human data, they see predictable improvements in the model's ability to accurately predict and simulate human behavior.

Related Insights

A 10x increase in compute may only yield a one-tier improvement in model performance. This appears inefficient but can be the difference between a useless "6-year-old" intelligence and a highly valuable "16-year-old" intelligence, unlocking entirely new economic applications.

Joon Sung Park's team chose simulation over personal agents because a useful assistant requires a deep, accurate model of its user's preferences and behaviors first. Understanding the person is a prerequisite for effective automation.

AI model capabilities follow a predictable, non-linear scaling law: increasing training compute by 10x roughly doubles a model's capabilities. This exponential relationship, rather than an incremental one, is what will drive underappreciated and disruptive advancements across many industries.

Today's AI boom is fueled by scaling computation, which is a known engineering challenge. The alternative, embedding nuanced, human-like inductive biases, is far harder as it requires a deep understanding of the problem space. This difficulty gap explains why massive models dominate AI development over more targeted, efficient ones—scaling is simply the more straightforward path.

A key surprise in AI development was the non-linear impact of scale. Sebastian Thrun noted that while AI trained on millions of documents is 'fine,' training it on hundreds of billions creates an 'unbelievably smart' system, shocking even its creators and demonstrating data volume as a primary driver of breakthroughs.

Professor Kyunghyun Cho highlights a key tension in AI research. High-fidelity predictive models (like OpenAI's Sora) are computationally regular and scalable on current hardware. However, human-like intelligence relies on abstract, high-level reasoning that skips unnecessary details, a more efficient but computationally challenging approach.

LLMs trained on online text often reflect what people say, not what they do. Simile bridges this 'say-do gap' by collecting real behavioral data and personal life stories through partners like Gallup. This grounds their agent simulations in reality, making them more predictive of actual behavior.

Like human experts, advanced AI models improve their answers the more time they spend on a problem. This 'inference scaling' means short evaluations may fail to capture a model's true capabilities, as performance continues to increase with more computation, making it difficult to establish a performance ceiling.

Simile's validation method involves collecting extensive data from real people, creating their "digital twins," and then testing if the twins can accurately predict how the real individuals behave in unseen experiments. This grounding in real-world data builds trust in the simulation's outputs.

Contrary to the narrative that model performance is plateauing, Demis Hassabis states that while returns from scaling are no longer exponential, they remain 'very substantial.' Frontier labs continue to see significant gains from increasing model size and compute, suggesting the current AI paradigm is not yet exhausted.