We scan new podcasts and send you the top 5 insights daily.
The API outputs—`no`, `score`, and `choice`—are intentionally designed as new concepts rather than mapping directly to existing types like booleans or integers. A `no` is a continuous probability, not a binary true/false. This forces developers to think differently about integrating probabilistic AI logic into code.
Making an API usable for an LLM is a novel design challenge, analogous to creating an ergonomic SDK for a human developer. It's not just about technical implementation; it requires a deep understanding of how the model "thinks," which is a difficult new research area.
Hype suggests JEV is a better, faster ChatGPT, but it's a fundamentally different tool. JEV is designed for machine-to-machine automation, outputting structured decisions and probabilities, not human-like text. It complements, rather than competes with, models like Claude or ChatGPT, and is not for direct human interface.
Instead of providing a 'seed' for deterministic outputs, Jev prioritizes robustness: ensuring similar inputs produce similar outputs. This is more critical for real-world software, which must handle slight variations gracefully. Strict determinism is a less important property that can be traded off for better cost and performance.
Unlike traditional software with deterministic outputs, generative AI systems require a new paradigm. Chip Huyen calls this "evaluation-driven development," where the focus shifts from writing fixed tests to building robust systems and guidelines for evaluating ambiguous, generative outputs.
When using an LLM to evaluate another AI's output, instruct it to return a binary score (e.g., True/False, Pass/Fail) instead of a numbered scale. Binary outputs are easier to align with human preferences and map directly to the binary decisions (e.g., ship or fix) that product teams ultimately make.
Traditional software relies on binary if-then statements. New judgment models like JEV fundamentally upgrade this by allowing those `if` conditions to understand "messy human context." This enables automation of complex processes like fraud detection, support routing, and lead scoring that previously required human interpretation of nuanced situations.
AI agents struggle to reliably differentiate between nuanced scores like '3 out of 5' versus '4 out of 5.' For effective self-correction in automated workflows, structure your evaluations (evals) as a series of unambiguous, binary pass/fail checks.
Jev is a classifier AI that makes probabilistic decisions based on predefined choices (a schema). Unlike LLMs which generate text conversationally, Jev provides structured, type-safe output, making it an "AI decision maker" rather than a chat agent that you "ask" questions.
Jev's output isn't a single definitive answer but a probability score for each possible choice (e.g., "80% confident this is a high-priority lead"). This structured, "type-safe" data allows developers to set thresholds and build complex, nuanced business logic directly in their code without parsing text.
The optimal way to use decision models like Jev is to break large problems into many small, independent questions. This contrasts with stuffing everything into a single LLM prompt. This decomposition makes each AI-driven step verifiable, measurable, and debuggable, leading to more reliable and maintainable software.