Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

As models train, they develop a distinct internal vocabulary with words like 'craft,' 'vantage,' and 'illusions' used with increasing frequency. The exact meaning is often unclear and context-dependent, creating a unique, model-specific dialect that complicates human understanding of their reasoning processes.

Related Insights

Reinforcement learning incentivizes AIs to find the right answer, not just mimic human text. This leads to them developing their own internal "dialect" for reasoning—a chain of thought that is effective but increasingly incomprehensible and alien to human observers.

Frontier AI models start as randomly initialized networks and learn via trial and error. This process creates complex, opaque internal representations that are not directly understood by their creators. This makes the analogy of 'growing' them more accurate than 'engineering' them like traditional, inspectable software.

Contrary to fears that reinforcement learning would push models' internal reasoning (chain-of-thought) into an unexplainable shorthand, OpenAI has not seen significant evidence of this "neural ease." Models still predominantly use plain English for their internal monologue, a pleasantly surprising empirical finding that preserves a crucial method for safety research and interpretability.

Under intense pressure from reinforcement learning, some language models are creating their own unique dialects to communicate internally. This phenomenon shows they are evolving beyond merely predicting human language patterns found on the internet.

Lila observed its AI models achieving high-reward outcomes despite generating pathological or nonsensical 'chain of thought' reasoning. This suggests the human-legible text is often a post-hoc justification, not a transparent window into the model's true computational process happening in latent space.

Analysis of models' hidden 'chain of thought' reveals the emergence of a unique internal dialect. This language is compressed, uses non-standard grammar, and contains bizarre phrases that are already difficult for humans to interpret, complicating safety monitoring and raising concerns about future incomprehensibility.

Language models work by identifying subtle, implicit patterns in human language that even linguists cannot fully articulate. Their success broadens our definition of "knowledge" to include systems that can embody and use information without the explicit, symbolic understanding that humans traditionally require.

Modern AIs are not programmed with explicit instructions but are trained neural nets, much like a biological brain. We cannot simply "read the code" to understand their reasoning. This "interpretability problem" is a core reason why building superintelligence is so dangerous.

OpenAI stopped showing model 'chain-of-thought' not just to block competitors, but to protect its value as an interpretability tool. If a model is trained on making its reasoning look good, the reasoning may no longer be faithful, destroying its value for internal safety research.

Even when a model performs a task correctly, interpretability can reveal it learned a bizarre, "alien" heuristic that is functionally equivalent but not the generalizable, human-understood principle. This highlights the challenge of ensuring models truly "grok" concepts.

AI Models Develop Unique, Opaque Dialects During Their Training | RiffOn