We scan new podcasts and send you the top 5 insights daily.
Unlike school's constant, direct feedback, professional life offers very little. This "wicked learning environment" makes improvement nearly impossible for roles like judges or managers. The key is not just to train individuals better, but to redesign professional systems to provide the feedback loops that drive learning.
Unlike traditional software where problems are solved by debugging code, improving AI systems is an organic process. Getting from an 80% effective prototype to a 99% production-ready system requires a new development loop focused on collecting user feedback and signals to retrain the model.
Jerry Tworek, a self-described "RL maximalist," found that scaling RL at OpenAI improved benchmarks but failed to solve real-world problems. The training data and evals were a closed loop, disconnected from the messy distribution of real user tasks, necessitating models that can learn at test time.
The "10,000 hours to mastery" concept is misunderstood. It works for domains with clear, repeating rules like chess, but not for "wicked" modern careers where rules change and reinvention is required. For most professionals, developing a broad range of skills is more valuable.
Intuition excels in areas like chess or boxing where we get immediate, repeated feedback. It fails in complex domains like choosing a charity or making social policy, where feedback is slow, noisy, or nonexistent. We mistakenly trust our intuition in these low-feedback environments where it's unreliable.
To accelerate growth for talented individuals, give them responsibility where their failure rate is between one-third and two-thirds. Most corporate roles are over-scaffolded with a near-zero chance of failure, which stifles learning. High potential for failure is a feature, not a bug.
Users often abandon AI when its first output is poor, akin to firing a new employee after their first attempt. Instead, train AI by providing clear, specific, behavior-based feedback repeatedly. It learns from reinforcement just like a human, but at a vastly accelerated rate.
Specialization thrives in "kind" environments like chess or golf, where rules are fixed and feedback is immediate. However, in "wicked" environments with unclear rules and delayed feedback—common in modern business—specialists struggle to adapt. Generalists, with broader experience, are better equipped for novel challenges.
Traditional training is ineffective for AI because models and best practices evolve too quickly. Companies like PricewaterhouseCoopers use dynamic "learning arenas"—like 'prompting parties'—where employees experiment and share discoveries in real-time. This creates a continuously adapting knowledge base that a static curriculum cannot match.
Demis Hassabis identifies a key obstacle for AGI. Unlike in math or games where answers can be verified, the messy real world lacks clear success metrics. This makes it difficult for AI systems to use self-improvement loops, limiting their ability to learn and adapt outside of highly structured domains.
We rigorously test software upgrades in a staging environment before going live, yet we expect humans to adopt new skills immediately after a training session. Employees need safe spaces to practice new behaviors, like communication, through repetition.