A model correctly deciding a customer needs a refund is easy. The engineering challenge is the production system that handles API calls, database updates, network failures, and rollbacks to turn that decision into a reliable outcome. This operational complexity is the true hurdle.
An agent may 'remember' an account is active, but the database could show it was suspended minutes ago. System architecture must distinguish between what the model believes and what the system knows to be true from its authoritative data sources, preventing actions based on stale information.
Don't rely on a prompt like 'never issue a large refund.' Instead, the application software should have a hardcoded, testable check. This 'AI proposes, software enforces' model creates a reliable safety net for business-critical invariants that cannot be violated.
Focusing on prompt injection is too narrow. A production AI's true vulnerabilities are systemic: influencing retrieved documents, poisoning agent memory, abusing tool permissions, or exploiting connected APIs. Security must be a system-wide, architectural concern.
Reducing AI costs is an engineering challenge, not just a procurement one. True optimization comes from architectural solutions like caching, request deduplication, routing simple tasks to smaller models, and, most importantly, deciding if a model call is even necessary.
