One AI Safety Contract, Multiple Enforcement Points: Why Prompts Are Not Enough
A new framework argues that AI agent safety cannot rely solely on prompts or instructions, since agents can bypass them without structural controls in place. The proposed approach centers on a single declared contract file, enforced at multiple mandatory checkpoints rather than through informal guidance. A local runner provides fast feedback and early refusals, but is considered optional and insufficient as a final enforcement layer. CI pipeline checks, made mandatory through branch protection policies, serve as the first non-optional gate to verify that only contract-compliant changes are merged. Runtime sandbox and capability boundaries form the second mandatory enforcement point, ensuring agents cannot reach unintended external effects even after passing CI.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in