AI agents spend most of their inference time on high-volume, repetitive tasks like embedding and reranking, not on the single generation step that users see. These 'underwater' tasks are best handled by small, specialized models, which dictates overall cost and latency.
Standard inference tooling is designed for one large model on many GPUs. Efficiently serving multiple small models requires the opposite architecture: packing many models onto a single GPU with fast switching to avoid paying for idle hardware, a fundamentally different infrastructure problem.
Despite potential cost savings, self-hosting is not always best. For low-volume or spiky traffic—under roughly 5 million requests or a total inference bill under $2,000 per month—the operational overhead outweighs the benefits, making hosted APIs the more economical option.
