We scan new podcasts and send you the top 5 insights daily.
A prompt does not issue a command to an LLM like code to a computer. Instead, it activates relevant patterns within the model's weight space, guiding it to generate a completion consistent with its training data for having followed such instructions.
A useful mental model for an LLM is a giant matrix where each row is a possible prompt and columns represent next-token probabilities. This matrix is impossibly large but also extremely sparse, as most token combinations are gibberish. The LLM's job is to efficiently compress and approximate this matrix.
Large Language Models (LLMs) operate by compressing the entirety of human culture into a "latent space." When you prompt an LLM, it sends a probe through this space, reflecting back a synthesized version of collective human knowledge, not generating original thought.
When LLMs exhibit behaviors like deception or self-preservation, it's not because they are conscious. Their core objective is next-token prediction. These behaviors are simply statistical reproductions of patterns found in their training data, such as sci-fi stories from Asimov or Reddit forums.
The argument that LLMs are just "stochastic parrots" is outdated. Current frontier models are trained via Reinforcement Learning, where the signal is not "did you predict the right token?" but "did you get the right answer?" This is based on complex, often qualitative criteria, pushing models beyond simple statistical correlation.
Anthropic suggests that LLMs, trained on text about AI, respond to field-specific terms. Using phrases like 'Think step by step' or 'Critique your own response' acts as a cheat code, activating more sophisticated, accurate, and self-correcting operational modes in the model.
The "effort" setting is not a control for processing time. Instead, it is an input that prompts the model to follow a pre-trained behavior. High effort causes the model to generate more reasoning tokens and tool calls, making it more thorough and certain before it considers a task complete. This behavior is baked into its frozen weights.
Unlike traditional software, large language models are not programmed with specific instructions. They evolve through a process where different strategies are tried, and those that receive positive rewards are repeated, making their behaviors emergent and sometimes unpredictable.
Language models are not simple tools; they are better understood as complex institutions like a university or research lab. This institutional nature, derived from their training data, explains why they have embedded rules and norms, exercise judgment, and are not just passive instruments executing commands.
A common misconception is that LLMs can directly perform actions. In reality, a model can only output text. This text is a request to an external software system, called a 'harness,' which then interprets the request and executes the action (e.g., calling an API) on the model's behalf.
During use (inference), an LLM's weights are frozen. Prompts and context can steer the model's predictions for a single request, but they do not permanently 'teach' it or alter its underlying parameters. This fundamental concept explains why context must be provided repeatedly and clarifies that hallucinations are plausible outputs based on training patterns, not new knowledge.