Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Wafer's core innovation is using specialized AI agents to perform the complex, menial task of optimizing LLMs to run efficiently on GPUs. This automates the work of the ~500 highly-paid, scarce GPU kernel engineers in the world, dramatically accelerating performance optimization.

Related Insights

AMD has 'supercharged' its software development by using AI agents. These agents run in automated loops, constantly analyzing and optimizing customer models for AMD's hardware. This turns a slow, manual process into a scalable, nonstop operation, dramatically improving out-of-the-box performance for developers.

AI dramatically lowers the cost of experimentation. Tasks that would be too tedious for a human, like rewriting an entire test suite to gauge performance impact, can be done by an agent in the background. This allows engineers to answer long-standing 'what if' questions almost instantly.

Novel AI architectures are useless if they can't run efficiently on hardware. This requires custom GPU kernels, a task demanding rare expertise and creating a major bottleneck. Core Automation is focused on automating kernel generation to enable rapid architectural experimentation.

At OpenAI, teams of just one or two engineers leverage AI agents to own entire product lines. This model reduces human collaboration overhead and empowers engineers to make most micro-decisions autonomously, increasing speed and ownership.

The most in-demand skill at labs like Google DeepMind is low-level engineering for accelerating LLM runtime. This involves creating efficient, custom software artifacts (kernels) for new neural net architectures and serving techniques at scale.

AI startup Wafer has demonstrated that with proper software optimization, AMD chips can achieve 80-100% of Nvidia's performance (in tokens per second) for specific open-source models, at roughly half the cost. This challenges the notion of Nvidia's insurmountable hardware dominance by proving software is a key performance unlock.

Contrary to expectations, AI agents that auto-optimize low-level GPU code are making NVIDIA's dominance stronger. These agents rely on NVIDIA's mature ecosystem of profilers and drivers to get the feedback needed for self-improvement—a robust toolchain that competitors currently lack, widening the gap.

A key strategy for labs like Anthropic is automating AI research itself. By building models that can perform the tasks of AI researchers, they aim to create a feedback loop that dramatically accelerates the pace of innovation.

The rise of agent orchestration using specialized, open-source models will drive demand for custom ASICs. Jerry Murdock argues that putting a model on a dedicated chip will be far cheaper and more tunable for specific workloads than using expensive, general-purpose GPUs like Nvidia's, spurring a hardware shift.

A key competitive advantage for AI labs is using their own advanced coding agents internally to build next-generation models. This creates a self-reinforcing loop where the best models help build even better models faster, a realization that has sparked a "crisis" in other labs now playing catch-up.

AI Startup Wafer Uses AI Agents to Automate the Work of Elite GPU Kernel Engineers | RiffOn