Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A key property of Recursive Language Models (RLMs) is their ability to decompose a large, complex task that is out-of-distribution (OOD) for the model. The harness breaks it into a series of smaller sub-problems, each of which is locally in-distribution, ensuring more reliable performance at each step.

Related Insights

A useful mental model for an LLM is a giant matrix where each row is a possible prompt and columns represent next-token probabilities. This matrix is impossibly large but also extremely sparse, as most token combinations are gibberish. The LLM's job is to efficiently compress and approximate this matrix.

In a 2018 interview, OpenAI's Greg Brockman described their foundational training method: ingesting thousands of books with the sole task of predicting the next word. This simple predictive objective was the key that unlocked complex, generalizable language understanding in their models.

Instead of building a single, monolithic AI agent that uses a vast, unstructured dataset, a more effective approach is to create multiple small, precise agents. Each agent is trained on a smaller, more controllable dataset specific to its task, which significantly reduces the risk of unpredictable interpretations and hallucinations.

The argument that LLMs are just "stochastic parrots" is outdated. Current frontier models are trained via Reinforcement Learning, where the signal is not "did you predict the right token?" but "did you get the right answer?" This is based on complex, often qualitative criteria, pushing models beyond simple statistical correlation.

A harness's design is an opinionated program that shapes how a model approaches a problem. A well-designed harness, like an RLM, can dramatically increase a model's generalization capabilities by providing a structural prior that helps it solve tasks more efficiently.

When asked to analyze 100 papers, LLMs often admit they didn't complete the task. This failure stems from outcome-based training, which prioritizes a plausible-looking final output over correctly following the required process, revealing a fundamental flaw in current training paradigms.

RLMs generalize effectively because they learn the abstract structure of a solution, which often remains consistent across seemingly different tasks. By training an RLM on one task, it can immediately solve another unrelated task if the underlying problem-solving 'program' is the same.

The idea that LLMs just predict the next statistical character is wrong. To accurately predict the next word in a physics paper, for instance, an AI must create an internal model of the physical world. This implies that scaling current architectures can lead to deeper understanding, not just pattern matching.

For unpredictable situations where a robot has no prior training data (e.g., a "gas leak" sign), multimodal LLMs can provide the necessary world knowledge to reason and act appropriately. This solves the long-standing robotics problem of how to handle the long tail of real-world scenarios.

LLMs are trained to produce high-probability, common information, making it hard to surface rare knowledge. The solution is to programmatically create prompts that combine unlikely concepts. This forces the model into an improbable state, compelling it to search the long tail of its knowledge base rather than relying on common associations.

RLMs Succeed by Turning OOD Problems into In-Distribution Sub-Tasks | RiffOn