We scan new podcasts and send you the top 5 insights daily.
Moving beyond the simple Linear Representation Hypothesis, models organize concepts within sparse mixtures of subspaces. The specific geometry of these "manifolds" (e.g., a circle for days of the week) encodes the relationships and valid operations between concepts, like chemistry emerging from the periodic table.
Naive model steering often fails because it cuts across a model's internal concept geometry, going "off-manifold" into meaningless space. By contrast, steering *along* the learned manifold—like tracing the circle from "Monday" to "Friday"—allows for smooth, effective control without degrading model performance into gibberish.
Human understanding is the ability to connect new information to a global, unified model of the universe. Until recently, AI models were isolated (e.g., a chess model). The major advance with large multimodal models is their ability to create a single, cohesive reality model, enabling true, generalizable understanding.
The field is moving beyond labeling concepts with sparse autoencoders. The new frontier is understanding the intricate geometric structures (manifolds) these concepts form in a model's latent space and how circuits transform them, providing a more unified, dynamic view.
Unlike classic theories based on simple equations, large AI models represent a new kind of scientific object. Rather than being mere predictive tools, they could be a novel form of explanation that we must learn to manipulate through new operations like distillation and merging, much like Mathematica made massive equations workable.
A common misconception is that Transformers are sequential models like RNNs. Fundamentally, they are permutation-equivariant and operate on sets of tokens. Sequence information is artificially injected via positional embeddings, making the architecture inherently flexible for non-linear data like 3D scenes or graphs.
Using a sparse autoencoder to identify active concepts, one can project a model's gradient update onto these concepts. This reveals what the model is learning (e.g., "pirate speak" vs. "arithmetic") and allows for selectively amplifying or suppressing specific learning directions.
Researchers modeling the 50 million neural connections in a fruit fly's brain found standard 3D spatial models were poor predictors. The most accurate model required a 64-dimensional framework, suggesting consciousness and cognition arise from a biological complexity far beyond our three-dimensional comprehension.
We can now prove that LLMs are not just correlating tokens but are developing sophisticated internal world models. Techniques like sparse autoencoders untangle the network's dense activations, revealing distinct, manipulable concepts like "Golden Gate Bridge." This conclusively demonstrates a deeper, conceptual understanding within the models.
While 'probes' require knowing what concept you're looking for, sparse autoencoders analyze a model's complex internal state (like white light) and automatically separate it into thousands of individual concepts (like a prism creating a rainbow). This can reveal concepts researchers hadn't thought to look for.
Human intelligence is multifaceted. While LLMs excel at linguistic intelligence, they lack spatial intelligence—the ability to understand, reason, and interact within a 3D world. This capability, crucial for tasks from robotics to scientific discovery, is the focus for the next wave of AI models.