We scan new podcasts and send you the top 5 insights daily.
Life has two modes: work (exploitation) and play (exploration). Genuine learning, for both humans and AI, happens exclusively during play—a low-stakes environment for experimentation. Work is simply applying what has already been learned, a concept mirrored in AI's exploration vs. exploitation trade-off.
DeepMind's core breakthrough was treating AI like a child, not a machine. Instead of programming complex strategies, they taught it to master tasks through simple games like Pong, giving it only one rule ('score go up is good') and allowing it to learn for itself through trial and error.
Contrary to the typical focus on efficiency, the most valuable discoveries with AI often come from unstructured exploration. Dedicate off-peak hours (e.g., 7 p.m. to 7 a.m.) to follow tangents and experiment creatively without the pressure of immediate productivity.
A child's seemingly chaotic learning process is analogous to the 'simulated annealing' algorithm from computer science. They perform a 'high-temperature search,' randomly exploring a wide range of possibilities. This contrasts with adults' more methodical 'low-temperature search,' which involves making small, incremental changes to existing beliefs.
To accelerate learning in AI development, start with a project that is personally interesting and fun, rather than one focused on monetization. An engaging, low-stakes goal, like an 'outrageous excuse' generator, maintains motivation and serves the primary purpose of rapid skill acquisition and experimentation.
Organizations fail when they push teams directly into using AI for business outcomes ("architect mode"). Instead, they must first provide dedicated time and resources for unstructured play ("sandbox mode"). This experimentation phase is essential for building the skills and comfort needed to apply AI effectively to strategic goals.
The brain circuits for play are not pruned after childhood; they persist because they are vital for adult adaptation. Biology doesn't waste resources. The continued existence of these circuits is proof that play is a fundamental, non-negotiable mechanism for learning and creativity throughout our entire lives.
Instead of only learning at test time, models should have a phase to retreat from live interaction and deeply integrate new information. This 'dreaming' allows them to experiment with their affordances and what they know, analogous to how humans consolidate memories.
A genuinely continual learner doesn't have separate training and testing phases. Instead, its life is a continuous process divided into two modes: an 'active' phase of interacting with new data and an 'offline' sleep phase for memory consolidation and self-improvement.
Play is not just for children or sports; it's a critical adult activity for exploring 'if-then' scenarios in a safe environment. This process of low-stakes contingency testing expands our mental catalog of potential outcomes, directly improving creativity and adaptability in high-stakes situations.
Research shows crows derive pleasure from using tools. We do this in our personal lives but approach work tech joylessly, aiming for basic competence. Instead, carve out time to "play" with new tools like AI. This childlike exploration increases enjoyment and mastery, reducing frustration and improving outcomes.