Home › Glossary › Safety › Agent Trust

Intermediate · Safety

Agent Trust

Visual diagram · (in preparation) · Math · (in preparation) · Worked example · 3 difficulty levels.

TL;DR. The calibrated confidence a system or person places in an agent, grounded in verified identity, bounded authority, and observed behaviour.

Technical Definition

The calibrated confidence a system or person places in an agent, grounded in verified identity, bounded authority, and observed behaviour.

How it works

Agent trust is earned and bounded, not assumed. It is built from verifiable identity, explicitly granted and narrow capabilities, evidence of past actions, and independent verification of outputs. Practical trust models are graduated: an agent may act autonomously in low-risk scopes, require review for medium-risk actions, and be blocked from irreversible ones. Crucially, trust in an agent must never be inherited by content the agent merely relays — trust attaches to the actor, not to the data it carries.

Related Concepts

  • Context Poisoning — An attack or accident in which false or hostile content enters an agent's context and corrupts its subsequent reasoning and actions.
  • Agent Identity — The verifiable answer to 'which agent is this?' — a stable, attestable identifier distinct from the human or service behind it.
  • Agent Action Evidence — Verifiable records proving what an agent did, under whose authority, with what inputs and result.
  • Agent Transparency — Making an agent's identity, authority, reasoning, and actions legible to the humans and systems it interacts with.