Governance rules are a form of code and can have bugs. Test them by having agents unfamiliar with the system rule on past decisions ("blind adjudication"). Also, have an agent intentionally try to find loopholes ("adversarial read") to ensure the rules are robust and unambiguous before they go live.
Instead of relying solely on role-based permissions, classify actions by their potential impact. Reversible actions within a domain can be automated (Green), those affecting other domains require owner consent (Amber), and irreversible actions like deletions or payments must require human approval (Red).
One AI project can exploit another's capabilities to perform an action it was denied. To prevent this, requests must carry only the requester's need, not its authority. The agent performing the action is always responsible for checking its own permissions, preventing it from being a "confused deputy."
In distributed systems, a component's physical location (e.g., a file in a directory) is not proof of its ownership. Responsibility must be explicitly recorded in a central registry. This "record over layout" principle prevents incorrect assumptions, especially for shared resources like scheduled jobs or configuration files.
Disagreements between projects often stem from misinterpreting information. By explicitly categorizing records as a verifiable "fact," a strategic "positioning choice," a "ruling," a "snapshot," or a "draft," the system can dissolve apparent contradictions before they escalate, such as when marketing copy is treated as a technical specification.
