Reinforcement Learning from Human Feedback (RLHF) forces models to be overly conservative to avoid obvious errors. This causes "mode collapse," where the model drops less common but valid possibilities, destroying its calibration and making it unreliable for programmatic decision-making that requires true confidence assessment.
An API that refuses to respond is fundamentally broken for software integration. This confuses "safety alignment" (appropriate for a consumer product like ChatGPT) with "capability alignment" (essential for a developer platform). For code, a stochastic failure is a critical bug, not a safety feature.
Contrary to the norm, TypeSafe AI avoids training on user data. They believe real-world data is heavily biased towards current use cases, which would cause the model to "fracture" and fail on future, unimagined applications. Their goal is a general cognitive core, not a model optimized for today's queries.
RLCD (Reinforcement Learning from Code Decisions) isn't just a new algorithm; it's a new 'task' or 'North Star' for AI development. It shifts the objective from RLHF's goal of 'pleasing humans' or RLVR's goal of 'winning benchmarks' to a new paradigm of creating reliable, verifiable outputs for software consumption.
Public benchmarks are seen as gamed and counterproductive because true intelligence has a 'je ne sais quoi' that leaderboards can't capture. The ultimate test is not a synthetic score but direct evaluation within a specific workflow. Long-term trust is built on reliability in production, not on winning benchmarks.
Instead of providing a 'seed' for deterministic outputs, Jev prioritizes robustness: ensuring similar inputs produce similar outputs. This is more critical for real-world software, which must handle slight variations gracefully. Strict determinism is a less important property that can be traded off for better cost and performance.
The optimal way to use decision models like Jev is to break large problems into many small, independent questions. This contrasts with stuffing everything into a single LLM prompt. This decomposition makes each AI-driven step verifiable, measurable, and debuggable, leading to more reliable and maintainable software.
Diogo Almeida claims that even with a billion-dollar investment, he would not engage in pre-training a new foundation model. He believes the most significant leverage and innovation comes from post-training techniques and superior data strategy, which can create more value than competing on raw compute for pre-training.
The API outputs—`no`, `score`, and `choice`—are intentionally designed as new concepts rather than mapping directly to existing types like booleans or integers. A `no` is a continuous probability, not a binary true/false. This forces developers to think differently about integrating probabilistic AI logic into code.
Offering separate models or modes for chat and reasoning is a flawed approach that fractures the model's underlying intelligence. Optimizing for conversational style (RLHF) inherently degrades calibration and logical consistency. The goal should be a single, smooth, reliable cognitive core, not specialized, conflicting versions.
Current AI can solve complex math problems yet fails to automate basic work, indicating a fundamental misapplication. The revolution is stalled because models are designed for human chat, not machine consumption. Creating economic value requires AI that integrates seamlessly into software, which has 'many nines' more automation potential.
Today's coding agents are architecturally limited by the KV cache, which forces an inefficient, append-only process within a single model. A better paradigm would free agents from this constraint, enabling proper software practices like state management, decomposition into sub-agents, and parallel execution for more powerful and scalable automation.
Before its viral launch, TypeSafe AI found that most potential customers didn't understand or see a need for its product. This challenges the conventional wisdom of finding product-market fit before a big launch. For truly novel technologies, a passionate developer community can create the market overnight, rendering prior feedback irrelevant.
