Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Ramin of Liquid AI argues that the current race to maximize model intelligence is flawed. A sustainable approach requires optimizing across three axes simultaneously: raw capability (intelligence), computational cost (efficiency), and deployment environment (substrate), moving beyond the data center to edge devices.

Related Insights

Significant opportunity exists in re-architecting how AI models work. Instead of building ever-larger single models, the focus is shifting to creating networks of smaller, specialized models that collaborate, which can drastically reduce the cost per token produced.

Breakthroughs like neural network "pruning" can reduce model size by 90% without losing accuracy, offering a 10x reduction in inference costs. This highlights that algorithmic innovation, not just acquiring more hardware, will be a key competitive vector in the AI race, enabling more output with less energy.

According to Liquid AI's CEO, the primary application of architectural research has become enabling efficiency—reducing cost, latency, and memory without sacrificing quality. The next major breakthroughs in AI *capability* are more likely to stem from new learning algorithms and data paradigms rather than architecture alone.

The era of using the most powerful AI model for every task is ending. Companies are now focused on the trade-off between quality, cost, and latency. The key question is no longer "Which model is best?" but "Which model is good enough for this task at the lowest price point?"

Models like Gemini 3 Flash show a key trend: making frontier intelligence faster, cheaper, and more efficient. The trajectory is for today's state-of-the-art models to become 10x cheaper within a year, enabling widespread, low-latency, and on-device deployment.

Liquid AI uses an automated system to discover neural architectures, avoiding human bias. Crucially, it bypasses misleading proxy metrics like perplexity by putting the target hardware in the loop and evaluating models directly on the customer's downstream tasks, optimizing for latency, memory, and quality.

Successful AI models will be small, specialized ones that run efficiently on consumer CPUs at the edge (laptops, phones). This leverages existing hardware (e.g., Apple's M-series chips) and avoids costly cloud GPUs, creating a strategic advantage for companies like Apple.

The key metric for winning the AI race is shifting from pure benchmark scores to efficiency. Perplexity's CEO argues that the company providing the most "token value per watt per user"—balancing accuracy, latency, cost, and intelligence—will ultimately dominate the market, making efficient intelligence the new goal.

The founder of Stormy AI focuses on building a company that benefits from, rather than competes with, improving foundation models. He avoids over-optimizing for current model limitations, ensuring his business becomes stronger, not obsolete, with every new release like GPT-5. This strategy is key to building a durable AI company.

The trend toward specialized AI models is driven by economics, not just performance. A single, monolithic model trained to be an expert in everything would be massive and prohibitively expensive to run continuously for a specific task. Specialization keeps models smaller and more cost-effective for scaled deployment.