We scan new podcasts and send you the top 5 insights daily.
Contrary to expectations, the most complex but verifiable fields like math and software engineering will likely be automated before subjective business functions. Verifiability provides clear training signals and objective success metrics for AI models, a luxury not present in areas like marketing or negotiations where "correctness" is fluid.
Andrej Karpathy's 'Software 2.0' framework posits that AI automates tasks that are easily *verifiable*. This explains the 'jagged frontier' of AI progress: fields like math and code, where correctness is verifiable, advance rapidly. In contrast, creative and strategic tasks, where success is subjective and hard to verify, lag significantly behind.
Silicon Valley is biased towards open-ended knowledge work like software engineering. However, a larger, often ignored opportunity for AI lies in automating the repeatable, deterministic business processes that power most of the non-tech economy, from customer support to operations.
Judgment Labs CEO Alex Shan argues that AI agents will first dominate domains with easily verifiable results, like coding, where a solution's correctness can be quickly checked. Progress will be slower in non-verifiable fields like law or complex drug discovery, where feedback loops are long and ambiguous.
While AI has mastered verifiable tasks with clear right answers, its future growth depends on human experts training models in subjective fields where 'good' is not easily defined. Companies are now sourcing professionals to act as 'verifiers' that teach AI nuanced, domain-specific judgment.
AI excels at solving problems with clear, verifiable answers, like advanced math, allowing for effective training. It struggles with complex societal issues like unemployment because there is no single, universally agreed-upon "correct" solution to train against, making it difficult to evaluate the AI's path.
AI will automate and replace jobs most rapidly in domains where its output can be objectively verified for correctness, like coding. In fields requiring subjective judgment with no single "right answer," such as creative or strategic roles, its impact will be augmentation, not outright replacement.
AI models improve dramatically in domains with objective feedback, like coding (unit tests) or science (lab results). Progress is slower in subjective fields like creative writing where feedback is opinion-based, explaining the uneven impact of AI across different types of knowledge work.
AI can generate vast amounts of content, but its value is limited by our ability to verify its accuracy. This is fast for visual outputs (images, UI) where our eyes instantly spot flaws, but slow and difficult for abstract domains like back-end code, math, or financial data, which require deep expertise to validate.
AI excels at generating code, making that task a commodity. The new high-value work for engineers is "verification”—ensuring the AI's output is not just bug-free, but also valuable to customers, aligned with business goals, and strategically sound.
Verifiability alone doesn't explain AI's rapid progress in math and coding. The key factor is 'grindability'—the ability to run thousands of parallel, containerized, and deterministic simulations. This allows for efficient credit assignment and learning, a luxury not available in domains like e-commerce or business strategy, which are constrained by real-world interactions and bot detectors.