We scan new podcasts and send you the top 5 insights daily.
Counter-intuitively, successful weather and climate AI models are not trained on long-term data. They are trained to predict only the next six hours autoregressively. This surprisingly generalizes to stable rollouts predicting weather patterns for hundreds or thousands of steps into the future, enabling long-term forecasting from short-term training.
It's surprising that AI models trained on general data can accurately predict rare events like hurricanes. The reason is that the physical world is "forgiving"; extreme phenomena are governed by strong physical structures and signatures that AI can learn effectively, even from a limited number of examples.
Unlike traditional neural networks which require fixed-resolution inputs (e.g., pixels), neural operators model data as continuous functions. This allows them to "zoom in" and make predictions at resolutions higher than the training data, a crucial capability for multi-scale physical phenomena like weather patterns.
AI's predictive power is based on identifying patterns in historical data. While effective when the future resembles the past, this makes it inherently unable to account for new inventions, crises, or paradigm shifts not represented in its training text. It predicts from old maps, not what will come next in a new world.
Traditional weather forecasting requires massive supercomputers, limiting access to large agencies. New AI models are tens of thousands of times faster and can run on a single consumer-grade GPU. This democratizes high-fidelity weather modeling for smaller agencies and nations, especially in the global south.
To build confidence in AI's ability to forecast the future, researchers are training "historical LLMs" on data ending in a specific year, like 1930. They then test the model's ability to predict text from a later period, like 1940. This process of historical validation helps calibrate and improve models predicting our own future.
Most AI weather models project the Earth onto a flat rectangle, causing simulations to become unstable and "blow up" over long periods. By incorporating the planet's spherical geometry using Fourier Neural Operators, models like ForecastNet remain stable for long-term climate rollouts, effectively becoming climate models.
While early AI development requires constant testing of new models, Conative.ai found they eventually reached a stable architecture. The focus then shifted from wholesale model replacement to fine-tuning existing layers with specific data, reducing the pressure to chase every new innovation.
Static data scraped from the web is becoming less central to AI training. The new frontier is "dynamic data," where models learn through trial-and-error in synthetic environments (like solving math problems), effectively creating their own training material via reinforcement learning.
A common forecasting error is to select the most likely outcome at each step. This creates an unrealistically 'normal' future. Realistic scenarios must instead sample from the distribution of possibilities, ensuring they include a plausible number of low-probability, high-impact events that shape the long-term trajectory.
When developing their Tario transformer model, Noetik discovered a key scaling behavior: larger, autoregressive models only outperform smaller ones when given a longer context window (i.e., seeing more tissue at once). This suggests that capturing broader spatial relationships is critical for learning complex biological patterns.