When an LLM fails, determine if it was a diligence issue (didn't try hard enough) or a capability issue (didn't know enough). This simple diagnostic framework helps decide whether to increase the model's effort level or upgrade to a larger, more knowledgeable model.
For difficult, multi-step tasks, a more capable LLM can reach a solution with fewer iterations than a smaller model. Despite a higher per-token price, this efficiency can lead to a lower total token count and a cheaper overall cost for the task, proving that cheaper-per-token isn't always cheaper-per-task.
The "effort" setting is not a control for processing time. Instead, it is an input that prompts the model to follow a pre-trained behavior. High effort causes the model to generate more reasoning tokens and tool calls, making it more thorough and certain before it considers a task complete. This behavior is baked into its frozen weights.
During use (inference), an LLM's weights are frozen. Prompts and context can steer the model's predictions for a single request, but they do not permanently 'teach' it or alter its underlying parameters. This fundamental concept explains why context must be provided repeatedly and clarifies that hallucinations are plausible outputs based on training patterns, not new knowledge.
