Sidekick part 4: safety as honestly labeled ceilings
Sidekick Part 4: safety as a stack of ceilings, each honestly labeled Most agent safety I've seen is a paragraph in the system prompt: "be careful with destructive commands." Small models ignore system prompts — we established that in Part 2. So Sidekick's safety model assumes the model will disobey and enforces the boundaries in code instead. This is Part 4: approvals, hard refusals, egress control, and the audit ledger — plus the ceilings, stated rather than hidden. Reads auto-run. Everything else sits in APPROVAL_TOOLS — writes, deletes, general shell — gated behind inline [y/N] prompts, wi
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in