Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

AI excels at tasks with clear verification (e.g., code review). The next breakthrough will be AI systems that can define "what good looks like" for subjective domains like law or management by creating their own evaluation frameworks, expanding AI's capabilities into new, complex areas.

Related Insights

AI excels where success is quantifiable (e.g., code generation). Its greatest challenge lies in subjective domains like mental health or education. Progress requires a messy, societal conversation to define 'success,' not just a developer-built technical leaderboard.

Verifying complex systems is bottlenecked by the human inability to specify all requirements. The future of software development is an interactive process where AI helps propose specifications (e.g., via test generation) and then uses a prover to formally verify them.

Andrej Karpathy's 'Software 2.0' framework posits that AI automates tasks that are easily *verifiable*. This explains the 'jagged frontier' of AI progress: fields like math and code, where correctness is verifiable, advance rapidly. In contrast, creative and strategic tasks, where success is subjective and hard to verify, lag significantly behind.

Judgment Labs CEO Alex Shan argues that AI agents will first dominate domains with easily verifiable results, like coding, where a solution's correctness can be quickly checked. Progress will be slower in non-verifiable fields like law or complex drug discovery, where feedback loops are long and ambiguous.

While AI has mastered verifiable tasks with clear right answers, its future growth depends on human experts training models in subjective fields where 'good' is not easily defined. Companies are now sourcing professionals to act as 'verifiers' that teach AI nuanced, domain-specific judgment.

Vinod Khosla highlights "auto-formalization" as a critical AI frontier. This technology converts ambiguous, human-written rules (e.g., legal code) into precise, machine-verifiable logic. This eliminates hallucinations, making AI reliable for mission-critical applications like tax law and medical diagnostics.

AI excels at generating code, making that task a commodity. The new high-value work for engineers is "verification”—ensuring the AI's output is not just bug-free, but also valuable to customers, aligned with business goals, and strategically sound.

Current benchmarks focus on whether code passes tests. The future of AI evaluation must assess qualitative, human-centric aspects like 'design taste,' code maintainability, and alignment with a team's specific coding style. These are hard to measure automatically and signal a shift toward more complex, human-in-the-loop or LLM-judged evaluation frameworks.

As AI masters content generation, it will handle the "blank page" problem. The crucial human task will then shift from creation to evaluation: defining what 'good' looks like, identifying AI failure modes, and building better verification systems to ensure outputs are trustworthy and useful.

For tasks where a simple right/wrong answer doesn't exist, verification is a major challenge. The solution is creating detailed rubrics with thousands of criteria, often developed with AI help. This provides a granular reward signal that allows models to climb the learning curve even in highly subjective domains.

The Next AI Frontier Is Systems That Build Their Own Verification for Ambiguous Tasks | RiffOn