When AI research firm METR tried to repeat its productivity study, 30-50% of developers declined to participate because they didn't want to forgo AI access. This selection bias makes establishing a true baseline for comparison nearly impossible, suggesting that measuring AI's true impact is becoming methodologically unfeasible as adoption grows.
An AI agent's verification loop only confirms its code satisfies the existing test suite; it does not validate that the overall approach was correct. An agent can pass every test while implementing a flawed architecture or introducing a vulnerability nobody thought to write a test for, widening the gap between a 'green checkmark' and human approval.
A novel vulnerability arises when an AI agent references a package name that doesn't exist. Malicious actors can register these hallucinated package names and upload malicious code. This creates a documented supply chain attack vector that requires specific checks beyond typical static analysis to mitigate.
A randomized controlled trial by AI evaluation non-profit METR showed experienced open-source developers were 19% slower when using AI tools. This contradicts the common narrative and the developers' own perception that they were 20% faster, highlighting a significant gap between perceived and measured productivity.
An analysis by Faro's AI found AI-assisted teams merged 98% more PRs, not by completing individual tasks faster, but by enabling developers to parallelize their workflow. Developers can kick off a task with an agent while simultaneously reviewing another human's work. This shows teams should optimize for throughput, not single-task velocity.
