Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

AI models excel in domains with discrete, quantifiable outcomes like coding or chess. However, they struggle with most knowledge work, which is often "unverifiable" and lacks a single correct answer for reinforcement learning models to train on. This distinction explains AI's current limitations in many professional roles.

Related Insights

Andrej Karpathy's 'Software 2.0' framework posits that AI automates tasks that are easily *verifiable*. This explains the 'jagged frontier' of AI progress: fields like math and code, where correctness is verifiable, advance rapidly. In contrast, creative and strategic tasks, where success is subjective and hard to verify, lag significantly behind.

AI performs poorly in areas where expertise is based on unwritten 'taste' or intuition rather than documented knowledge. If the correct approach doesn't exist in training data or isn't explicitly provided by human trainers, models will inevitably struggle with that particular problem.

Unlike coding, where context is centralized (IDE, repo) and output is testable, general knowledge work is scattered across apps. AI struggles to synthesize this fragmented context, and it's hard to objectively verify the quality of its output (e.g., a strategy memo), limiting agent effectiveness.

AI thrives in domains with fixed, written rules and searchable histories, like programming. In ambiguous areas like organizational conflict or political negotiation, where context is unwritten and lives in people's heads, its performance plummets. Its confident output masks this unreliability, posing a danger to decision-makers.

While AI has mastered verifiable tasks with clear right answers, its future growth depends on human experts training models in subjective fields where 'good' is not easily defined. Companies are now sourcing professionals to act as 'verifiers' that teach AI nuanced, domain-specific judgment.

Contrary to expectations, the most complex but verifiable fields like math and software engineering will likely be automated before subjective business functions. Verifiability provides clear training signals and objective success metrics for AI models, a luxury not present in areas like marketing or negotiations where "correctness" is fluid.

AI excels at solving problems with clear, verifiable answers, like advanced math, allowing for effective training. It struggles with complex societal issues like unemployment because there is no single, universally agreed-upon "correct" solution to train against, making it difficult to evaluate the AI's path.

Unlike coding, most real-world tasks lack training data that represents the task's actual execution. AIs are trained on descriptions and commentary, not performance data, akin to learning chess from analysis rather than gameplay. This severely limits their practical abilities in most domains.

AI will automate and replace jobs most rapidly in domains where its output can be objectively verified for correctness, like coding. In fields requiring subjective judgment with no single "right answer," such as creative or strategic roles, its impact will be augmentation, not outright replacement.

Demis Hassabis identifies a key obstacle for AGI. Unlike in math or games where answers can be verified, the messy real world lacks clear success metrics. This makes it difficult for AI systems to use self-improvement loops, limiting their ability to learn and adapt outside of highly structured domains.