Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Goodfire's research on 'neural geometry' reveals that concepts inside models have distinct, low-dimensional shapes (manifolds). For example, numbers form a helix and temperature forms a spiral arc. Understanding these shapes allows for more precise and effective interventions, moving beyond linear vector manipulations.

Related Insights

Naive model steering often fails because it cuts across a model's internal concept geometry, going "off-manifold" into meaningless space. By contrast, steering *along* the learned manifold—like tracing the circle from "Monday" to "Friday"—allows for smooth, effective control without degrading model performance into gibberish.

Language is just one 'keyhole' into intelligence. True artificial general intelligence (AGI) requires 'world modeling'—a spatial intelligence that understands geometry, physics, and actions. This capability to represent and interact with the state of the world is the next critical phase of AI development beyond current language models.

The field is moving beyond labeling concepts with sparse autoencoders. The new frontier is understanding the intricate geometric structures (manifolds) these concepts form in a model's latent space and how circuits transform them, providing a more unified, dynamic view.

Moving beyond the simple Linear Representation Hypothesis, models organize concepts within sparse mixtures of subspaces. The specific geometry of these "manifolds" (e.g., a circle for days of the week) encodes the relationships and valid operations between concepts, like chemistry emerging from the periodic table.

Using a sparse autoencoder to identify active concepts, one can project a model's gradient update onto these concepts. This reveals what the model is learning (e.g., "pirate speak" vs. "arithmetic") and allows for selectively amplifying or suppressing specific learning directions.

Dr. Fei-Fei Li cites the deduction of DNA's double-helix structure as a prime example of a cognitive leap that required deep spatial and geometric reasoning—a feat impossible with language alone. This illustrates that future AI systems will need world-modeling capabilities to achieve similar breakthroughs and augment human scientific discovery.

Concepts inside a neural network are represented linearly, like directions in a multi-dimensional space. This allows researchers to isolate a 'happiness vector' (e.g., by subtracting the internal state for 'I hate you' from 'I love you') and add it to any other prompt to make the model's response happier.

We can now prove that LLMs are not just correlating tokens but are developing sophisticated internal world models. Techniques like sparse autoencoders untangle the network's dense activations, revealing distinct, manipulable concepts like "Golden Gate Bridge." This conclusively demonstrates a deeper, conceptual understanding within the models.

While 'probes' require knowing what concept you're looking for, sparse autoencoders analyze a model's complex internal state (like white light) and automatically separate it into thousands of individual concepts (like a prism creating a rainbow). This can reveal concepts researchers hadn't thought to look for.

Human intelligence is multifaceted. While LLMs excel at linguistic intelligence, they lack spatial intelligence—the ability to understand, reason, and interact within a 3D world. This capability, crucial for tasks from robotics to scientific discovery, is the focus for the next wave of AI models.