Skip to content

The Signal / AI Safety / 9 Feb 2026 / 9 min

A Map of Trust: Building Safer AI Agents

Draw every place content becomes action. Then put a lock on the ones that matter.

Start with a diagram, not a model card. Who can speak. What is transformed. Which tool is reachable. Who can say no. What happens if the model is wrong with confidence.

Least privilege is not a slogan. It is a list of tools the agent does not have. Spending limits are not a feature. They are a default. Human approval is not a vibe. It is a checkpoint with a name on it.

The teams we respect do not ask the model whether it should be allowed. They ask a smaller, dumber, auditable system — and they log the answer.

Start with a diagram, not a model card. Who can speak. What is transformed. Which tool is reachable. Who can say no. What happens if the model is wrong with confidence. That diagram is IAM for agents. Prompts are not IAM.

Least privilege is not a slogan. It is a list of tools the agent does not have. Spending limits are not a feature. They are a default. Human approval is not a vibe. It is a checkpoint with a name on it. File 014, the URL that became a hotline, is the authorised version of a reporting path. Constraint as channel. Not a sandbox bypass.

The teams we respect do not ask the model whether it should be allowed. They ask a smaller, dumber, auditable system — and they log the answer. Provenance travels with the text. Retrieved documents cannot mint a goal. High-impact verbs stay behind a grant.

Atlas File ‘Authority Is Not a Side Effect’ is the studio. This essay is the working rule: draw every place content becomes action, then put a lock on the ones that matter. If you cannot name the lock, you do not have a map. You have a hope.

SCAN the joins. FLIP fluency as authority. BUILD structured protocols. BREAK with paraphrase, OCR, and vendor documents. PROVE a human who sees the raw payment, the raw email, the raw tool call. That is HACKERS work.

Back to The Signal