/
© 2026 RiffOn. All rights reserved.

Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

  1. Invest Like the Best with Patrick O'Shaughnessy
  2. Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy · Aug 25, 2026

SAIL Research founder Neil Movva is building a "token factory" to make AI 10x cheaper by optimizing for throughput over latency.

NVIDIA's "Speed of Light" Ethos Drives Its Dominance in Hardware Optimization

A core cultural tenet at NVIDIA is chasing the "speed of light"—pushing hardware to its absolute theoretical performance limit. This relentless focus on squeezing every ounce of capability out of their silicon is deeply ingrained in their engineering culture and is a key driver of their market leadership.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

NVIDIA Strategically Allocates Chips to Foster a Competitive Customer Ecosystem

NVIDIA doesn't simply sell its scarce chips to the highest bidder. It strategically allocates them to cultivate a diverse ecosystem of cloud providers and customers. This prevents any single customer from becoming too powerful and ensures healthy competition among its buyers, which ultimately drives more demand for NVIDIA's hardware.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

Cybersecurity is Becoming a "Proof-of-Work" Game Measured in AI Compute Spend

The security of software is increasingly determined by how much AI compute was spent trying to break it. Companies use diverse AI agents to autonomously attack their own code. The amount of money spent on APIs from labs like Anthropic to find vulnerabilities is becoming the best indicator of a system's robustness.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

GPUs Fundamentally Trade Throughput for Latency, Creating an Untapped Market

Optimizing a GPU for low-latency (fast, individual responses) inherently sacrifices its peak throughput (total work done over time). The market's focus on chatbots has over-indexed on latency, creating a major opportunity for companies that build systems optimized for high-throughput, background tasks.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

A "Scavenger Strategy" Unlocks Cheaper AI Compute By Using Unwanted Chips and Power

To achieve radical cost reduction, the strategy is to "scavenge" what others won't use: less popular chips (non-NVIDIA), stranded power from intermittent renewables, and small, low-reliability data centers. This arbitrage approach avoids competing for premium resources with deep-pocketed frontier labs.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

Low-Reliability (95% Uptime) Data Centers Are the Future of Cheap AI Inference

The need for high-availability data centers is an assumption from the training and real-time era. For asynchronous background agents, a distributed fleet of small, cheap data centers with 95% uptime is viable. Failures are handled by a robust control plane that reroutes work, trading P99 latency for unbeatable economics.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

"Latent Distillation" Ensures Open Source Will Always Catch Up to Closed Models

Preventing knowledge transfer from frontier models to open source is impossible. As AI-generated content (code, text) populates the internet, that data becomes part of the training set for the next generation of open-source models. This "latent distillation" ensures a constant diffusion of capabilities.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

AI's Next Data Frontier Is Self-Improvement on Verifiable Tasks, Not More Internet Data

The "one-time subsidy" of high-quality internet text is largely exhausted. The future of data for training models lies in creating reinforcement learning (RL) "gyms" where agents work on hard, verifiable problems (e.g., coding, math). The environment and the agent's progress become the new data source for recursive self-improvement.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

The Future of AI Is Asynchronous Agents, Making Latency Irrelevant

The dominant AI use case will shift from real-time, human-in-the-loop chatbots to long-running background agents. For these agents, which work for hours or days, an extra few seconds of latency is meaningless, unlocking massive cost-saving opportunities by prioritizing throughput over speed.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

Software Optimization Unlocks a Major Arbitrage in Non-NVIDIA AI Chips

The market undervalues chips from vendors like AMD because their software stack and kernel libraries are less mature than NVIDIA's CUDA. A team with deep expertise in low-level software and kernel optimization can extract significantly more performance from these chips, creating a powerful arbitrage opportunity by buying them at a discount.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

Specialized AI Chips Will Be Hybridized With GPUs, Not Replace Them

Chips like Cerebras and Grok, which excel at fast memory access (SRAM), are best suited for the compute-bound parts of a transformer model (the MLP). They will likely be paired with traditional GPUs, which are better at handling the memory-capacity demands of the attention mechanism and KV Cache, creating powerful hybrid systems.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago

The True Unlock for Proactive AI is Spending Tokens Without a Guaranteed Return

To create truly proactive, helpful AI assistants (like a better Siri), systems must be willing to spend vast amounts of cheap compute "speculatively" in the background without a direct user prompt or a guaranteed return on that spend. This is only possible when the cost of a token approaches zero.

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488] thumbnail

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·a month ago