We scan new podcasts and send you the top 5 insights daily.
Influential papers like RLMs, SWE-Bench, and Quiet-STaR were initially dismissed by many as simple or pointless. This initial negative reaction is often a signal of a good idea, as it indicates a departure from mainstream thinking that can unlock entirely new research avenues.
Citing Leopold Ashenbrenner's essay, the hosts argue that AI progress isn't linear. It relies on "unhovelers"—fundamental scientific discoveries like new attention mechanisms that unlock massive, non-linear gains, defying simple extrapolation of current trends.
When OpenAI started, the AI research community measured progress via peer-reviewed papers. OpenAI's contrarian move was to pour millions into GPUs and large-scale engineering aimed at tangible results, a strategy criticized by academics but which ultimately led to their breakthrough.
To pioneer neural machine translation, Prof. Kyunghyun Cho and his team deliberately limited their review of past research. They believed reading too much would impose false constraints from outdated contexts, preventing them from developing a system from scratch with fresh thinking.
Contrary to the "bitter lesson" narrative that scale is all that matters, novel ideas remain a critical driver of AI progress. The field is not yet experiencing diminishing returns on new concepts; game-changing ideas are still being invented and are essential for making scaling effective in the first place.
Current AIs are trained on the established, consensus-driven scientific literature. The real breakthrough will occur when AI is trained on the 'trash can corpus'—all the ideas and papers that were rejected, laughed at, and dismissed by the orthodoxy. This is where undiscovered alpha lies.
Socher reveals his pioneering 2018 paper on a unified NLP model, which heavily influenced the first GPT paper, was harshly rejected by academic reviewers. They deemed the concept of a single network for multiple tasks 'unfathomable,' halting his team's progress and delaying the field's advancement.
Great ideas like deep learning were not immediately recognized. Their value emerged over time as others built upon them. This suggests an idea's fruitfulness is a product of its context and cultural adoption, not just its isolated brilliance, making it difficult for an AI to evaluate its ultimate impact.
Despite the resource gap with industry, academia excels at fostering contrarian research. Stefano Ermon points to diffusion models, Flash Attention, and DPO—all with academic origins—as proof that this environment enables fundamental breakthroughs that industry might overlook.
Cohere's CEO believes if Google had hidden the Transformer paper, another team would have created it within 18 months. Key ideas were already circulating in the research community, making the discovery a matter of synthesis whose time had come, rather than a singular stroke of genius.
The foundational concept for modern LLMs, the attention mechanism, originated from an intern, Dima Badanao, in Yoshua Bengio's lab. The idea was so brilliant that its potential for success was immediately apparent upon explanation, before it was even coded.