Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

OpenAI treats its finite compute resources like capital. Instead of centralizing allocation, leadership gives individual product teams a fixed compute "budget." This forces teams on the ground to make difficult trade-offs and find creative efficiencies in their stack, ultimately unlocking more innovation and value than a top-down management approach would allow.

Related Insights

Unlike traditional software, OpenAI's growth is limited by a zero-sum resource: GPUs. This physical constraint creates a constant, painful trade-off between serving existing users, launching new features, and funding research, making GPU allocation a central strategic challenge.

Unlike compute-rich giants, AppLovin's bootstrapped culture enforces extreme efficiency in its AI infrastructure. Engineers don't have unlimited GPUs, forcing them to optimize code and models for cost and performance. This constraint-driven approach leads to significant cost savings and a lean operational model.

At OpenAI, teams of just one or two engineers leverage AI agents to own entire product lines. This model reduces human collaboration overhead and empowers engineers to make most micro-decisions autonomously, increasing speed and ownership.

To maximize optionality, OpenAI evolved from relying on a single cloud provider and chipmaker to a multi-faceted "Rubik's Cube" approach. This involves using multiple CSPs (Oracle, GCP, AWS) and chip providers (Nvidia, AMD) to ensure access to frontier technology while converting capital expenditures into operating expenses through partners.

Instead of managing compute as a scarce resource, Sam Altman's primary focus has become expanding the total supply. His goal is to create compute abundance, moving from a mindset of internal trade-offs to one where the main challenge is finding new ways to use more power.

OpenAI operates with a "truly bottoms-up" structure because it's impossible to create rigid long-term plans when model capabilities are advancing unpredictably. They aim fuzzily at a 1-year+ horizon but rely on empirical, rapid experimentation for short-term product development, embracing the uncertainty.

For a lean research team, the primary job isn't just building models but acting as investors allocating a scarce resource: compute. This capital allocator mindset focuses the team on placing bets on the most promising ideas and architectures, rather than spreading resources thin.

As AI agents run for longer periods, the primary decision is no longer just about engineering time but about allocating expensive compute resources. The product manager's role shifts to deciding which tasks are valuable enough to spend significant AI compute budget on, a decision made during the spec and planning phase.

Baidu forgoes rigid policies allocating AI compute 'tokens' to employees based on seniority or title. The CFO argues the unit cost of compute drops so fast that such policies become obsolete in weeks. They prefer empowering talent with ample resources, trusting them to prioritize tasks efficiently in a nimble environment.

To solve the challenge of budgeting for AI, Andrew MacDonald proposes a novel approach: merge the headcount and compute budgets into a single pool. This forces leaders to make direct trade-offs between hiring more engineers and spending on AI models, ensuring they allocate capital to the highest ROI activities.

OpenAI Decentralizes Compute Allocation to Force Product Team Efficiency | RiffOn