We scan new podcasts and send you the top 5 insights daily.
During use (inference), an LLM's weights are frozen. Prompts and context can steer the model's predictions for a single request, but they do not permanently 'teach' it or alter its underlying parameters. This fundamental concept explains why context must be provided repeatedly and clarifies that hallucinations are plausible outputs based on training patterns, not new knowledge.
A useful mental model for an LLM is a giant matrix where each row is a possible prompt and columns represent next-token probabilities. This matrix is impossibly large but also extremely sparse, as most token combinations are gibberish. The LLM's job is to efficiently compress and approximate this matrix.
Large Language Models (LLMs) operate by compressing the entirety of human culture into a "latent space." When you prompt an LLM, it sends a probe through this space, reflecting back a synthesized version of collective human knowledge, not generating original thought.
When LLMs exhibit behaviors like deception or self-preservation, it's not because they are conscious. Their core objective is next-token prediction. These behaviors are simply statistical reproductions of patterns found in their training data, such as sci-fi stories from Asimov or Reddit forums.
Karpathy identifies a key missing piece for continual learning in AI: an equivalent to sleep. Humans seem to use sleep to distill the day's experiences (their "context window") into the compressed weights of the brain. LLMs lack this distillation phase, forcing them to restart from a fixed state in every new session.
When an LLM is shown few-shot examples of a new task, it is performing Bayesian updating. With each example provided in the prompt, its belief (posterior probability) about the correct next token shifts, allowing it to "learn" a new pattern on the fly without changing its weights.
Rich Sutton argues that LLM-based assistants are not "experiential learners." Despite in-context learning, their fundamental weights never change after deployment. This prevents them from truly adapting, forming new concepts, or updating their core understanding based on user interactions.
The "memory" feature in today's LLMs is a convenience that saves users from re-pasting context. It is far from human memory, which abstracts concepts and builds pattern recognition. The true unlock will be when AI develops intuitive judgment from past "experiences" and data, a much longer-term challenge.
A fundamental misunderstanding is that AI learns from each interaction. It doesn't. Models are trained, but each new prompt is a fresh start, like 'Groundhog Day.' They operate on syntactic patterns without building semantic understanding or memory, which explains their inconsistent responses.
The "effort" setting is not a control for processing time. Instead, it is an input that prompts the model to follow a pre-trained behavior. High effort causes the model to generate more reasoning tokens and tool calls, making it more thorough and certain before it considers a task complete. This behavior is baked into its frozen weights.
A key gap between AI and human intelligence is the lack of experiential learning. Unlike a human who improves on a job over time, an LLM is stateless. It doesn't truly learn from interactions; it's the same static model for every user, which is a major barrier to AGI.