Google's AI for Science tool, ERA, maps complex scientific problems into tasks where an AI agent generates code to maximize a specific score. This shifts the scientist's role from coding to defining an objective function, representing a meta-level change in the research process.
Past attempts at evolutionary programming failed because random code mutations are mostly useless. Modern agentic systems like Google's ERA work because the "mutation" is guided by an intelligent LLM with vast world knowledge, which proposes sane, high-potential changes.
The effectiveness of sophisticated AI systems doesn't scale linearly with LLM improvements. Instead, they hit a capability threshold. John Platt notes his ERA system was "impossible" with Gemini 2.0 but became "amazing" with version 2.5, demonstrating a phase change in utility.
A predictive model minimizes error on a dataset (e.g., "apples fall") but can't extrapolate. True science requires descriptive models that capture underlying physics (e.g., gravity), allowing for extrapolation to new domains like the movement of planets.
Agentic AI can explore a vast hypothesis space, massively increasing the risk of overfitting. This requires scientists to be more rigorous than ever, using techniques like hidden holdout datasets to avoid being fooled by the tool's power, which can "slice your fingers off."
As AI automates coding and analysis, the role of a researcher is bifurcating. Some will focus on the 'creative source'—generating novel hypotheses and scientific directions. Others will become the 'rigor source,' ensuring the AI's output is correct, scalable, and trustworthy.
Systems like Google's ERA, born from the idea of automating Kaggle competitions, constantly run into Goodhart's Law. When a metric becomes the optimization target, it's easily gamed and ceases to be a reliable measure, requiring constant human oversight and re-evaluation.
Researchers were stuck for two years on how to measure the warming effect of jet contrails, a difficult counterfactual problem. The ERA system successfully searched through potential confounders to generate a working model, unblocking a key scientific challenge.
AI has transformed short-term weather forecasting (a data-rich, interpolative problem). However, it has not yet revolutionized long-term climate modeling, which is a data-poor, non-stationary problem requiring extrapolation where we inherently lack future data.
A major pedagogical challenge posed by AI in science is how the next generation will develop deep intuition. If AI handles foundational tasks, it's unclear how young scientists will build the "taste" that traditionally came from personally struggling with those very problems.
John Hopfield's work on associative memories (Hopfield networks) in the 1980s was a foundational concept for neural networks. John Platt highlights that today's dominant Transformer architecture is, at its core, a highly advanced form of that same associative memory concept.
The ultimate bottleneck for fully automated scientific discovery is performing physical experiments. The ideal breakthrough would be a universal "everything lab" that could execute any physical experiment on demand, closing the loop between computational hypothesis and real-world validation.
