Influential papers like RLMs, SWE-Bench, and Quiet-STaR were initially dismissed by many as simple or pointless. This initial negative reaction is often a signal of a good idea, as it indicates a departure from mainstream thinking that can unlock entirely new research avenues.
Models like JEV and concepts like RLMs and loop transformers signal a shift away from the dominant autoregressive decoder architecture. This opens up a new design space, allowing researchers to create models that trade off capabilities for benefits like extremely low inference latency.
RLMs generalize effectively because they learn the abstract structure of a solution, which often remains consistent across seemingly different tasks. By training an RLM on one task, it can immediately solve another unrelated task if the underlying problem-solving 'program' is the same.
A key property of Recursive Language Models (RLMs) is their ability to decompose a large, complex task that is out-of-distribution (OOD) for the model. The harness breaks it into a series of smaller sub-problems, each of which is locally in-distribution, ensuring more reliable performance at each step.
Deploying a large number of agents is the easy part. The difficult, unsolved problem is training the swarm to converge on a correct answer efficiently without generating excessive, useless output ('slop'). This is the non-trivial engineering feat behind successes like OpenAI's math proofs.
PhD students can't compete with industry labs on resources. Their unique advantage is the freedom to pursue non-obvious, big-bet research that industry might deem trivial or without immediate application. These unconventional bets are often the source of breakthrough ideas like SWE-bench or RLMs.
A harness's design is an opinionated program that shapes how a model approaches a problem. A well-designed harness, like an RLM, can dramatically increase a model's generalization capabilities by providing a structural prior that helps it solve tasks more efficiently.
While AI-generated solutions dominate GPU optimization leaderboards, the most stable and performant kernels are created by experts who actively guide and prompt the AI. This human-in-the-loop approach demonstrates that deep domain knowledge is still crucial for steering AI towards practical, production-ready solutions.
Despite different branding, popular agent harnesses like Claude Code, Codex, and Pi share the same fundamental logic: a loop that appends a trajectory to a prompt. For powerful models like GPT-4, the choice between them is irrelevant, suggesting the need for entirely new paradigms like RLMs.
Models like GPT-4 are disproportionately good at specific tasks like coding but can't transfer that abstract problem-solving skill to other domains, unlike a human expert. A key research frontier is bridging this gap to create more versatile and consistently capable AI systems.
