A Stanford study found that the vast majority of queries sent to powerful frontier models don't require their advanced capabilities. These tasks could be handled by smaller, faster, and more private local models at virtually no cost, revealing a massive inefficiency in the current API-centric approach.
Hugging Face's CEO dismisses the controversy around distillation, framing it as a widespread technique that offers only a marginal boost. It doesn't determine a model's fundamental quality—'if you suck, you suck with or without distillation'—and he questions the merit of 'unfair competition' claims from dominant, trillion-dollar companies.
Relying on a single frontier model is risky and inefficient. The next phase of AI will involve intelligently routing queries to the most appropriate model—be it cheaper, faster, or local. This will redistribute value from a few dominant labs to a long tail of specialized models, maturing the ecosystem.
Open source's safety extends beyond transparency. Its decentralized nature makes it less likely to be used for dangerous, large-scale projects like cyberweapons, which historically emerge from secretive, well-funded, closed-source efforts. The community naturally steers toward solving different, less risky problems.
With open-weight models, the user has full control, transparency, and access, mitigating risks of bias or manipulation from the creator. This is fundamentally different from using a foreign-hosted API, where you send them your data and they control access, making provenance a critical security concern.
Hugging Face's CEO argues that regulators' caution towards new models isn't surprising. Frontier labs spent years marketing their own models (like GPT-2) as dangerously powerful, which naturally led governments to take a more hands-on, safety-first approach to their deployment.
