We scan new podcasts and send you the top 5 insights daily.
Warp uses "human interactions per PR" as a core metric to quantify efficiency. This metric aggregates all human touchpoints—prompts, steering, PR comments, corrections—with the goal of minimizing them to increase agent autonomy and improve overall throughput.
Superhuman adopted AI coding tools using a three-quarter plan: 1) Unrestricted experimentation with centralized budget approval. 2) Analysis and measurement using self-reported PR labels. 3) Observing a sustained increase in engineering throughput from 4 to 6 PRs per engineer per week.
Front's research reveals a hidden "coordination tax" where teams spend the majority of their time on operational tasks like managing handoffs and re-explaining context. This 3:1 ratio of coordination-to-problem-solving cripples efficiency, even if traditional metrics like response time look good.
Intercom's CTO set a goal to 2x R&D throughput, using pull requests as a simple, albeit crude, metric. In a high-trust environment, this focused the team on adopting AI tools to increase output, leading to measurable success.
Traditional software development processes, like peer code reviews, were built for a cadence of 10-15 PRs per month. When AI agents enable a 10x increase in output, the human team becomes the bottleneck, forcing a shift towards AI-driven review and validation.
Gusto's "Cofounder" team achieved a median PR review time of just nine minutes, facilitated by a constant "PermaZoom" room where reviews could be requested and conducted instantly. This proves that ultra-fast human feedback loops, not just AI code generation, are the true enabler of rapid development.
With AI agents autonomously generating pull requests, the primary constraint in software development is no longer writing code but the human capacity to review it. Companies like Block are seeing PRs per engineer increase massively, creating a new challenge for engineering managers to solve.
In an agent-driven workflow, human review becomes the primary bottleneck. By moving reviews to after the merge, the team prioritizes agent throughput and treats human attention as a scarce resource for high-level guidance, not gatekeeping individual pull requests.
An analysis by Faro's AI found AI-assisted teams merged 98% more PRs, not by completing individual tasks faster, but by enabling developers to parallelize their workflow. Developers can kick off a task with an agent while simultaneously reviewing another human's work. This shows teams should optimize for throughput, not single-task velocity.
Warp's data shows that while AI can generate a PR in 35 minutes, the wait for a human review takes 3.5 hours. This demonstrates that even in highly automated development environments, human review processes remain the most significant drag on velocity.
Stripe's internal 'Minions' are AI agents that handle the entire coding workflow from prompt to PR. The key success metric is the percentage of PRs created in 'one shot'—without any human iteration—prioritizing true automation over mere assistance.