OpenAI has abandoned traditional benchmark score charts for its new GPT-6 models. Instead, they exclusively use graphs plotting performance against cost, forcing developers to evaluate models based on economic value and task-specific efficiency rather than just raw intelligence scores.
User preference for AI models is driven by more than just benchmarks. The 'personality' and interaction style of a model, like Opus 5.5's collaborative feel, directly impacts its perceived value and usability, making it a critical, and often overlooked, component of the user experience.
The cost for a given level of AI performance is falling at an unprecedented rate of 47% per quarter, according to Epic AI Research. This drop is multiples faster than Moore's Law, DNA sequencing, or electricity, unlocking previously uneconomical use cases like large-scale agent swarms.
A strategic divide is emerging: OpenAI (Sol/Luna) is optimizing for cheap, fast, iterative 'daily driver' models for frequent interaction. In contrast, Anthropic (Opus 5.5) targets high-performance models for ambitious coding and visual projects where users pay more for superior output.
Superior model performance alone no longer guarantees developer adoption. The surrounding 'harness'—the ecosystem of tools, stored context, and established workflows—creates significant inertia. Users now often stick with a 'good enough' model within their preferred ecosystem rather than migrate for incremental gains.
The current AI development strategy of 'pacing the frontier' focuses on refining existing model tiers rather than rushing to the next major capability jump. This strategy prioritizes cost reduction, efficiency, and fixing model flaws, leading to broader, more practical adoption over raw power.
