How Human Approval Works in AI Agents
Design approval gates that preserve speed while keeping authority with people.
Reviewed 2026-09-30
Human approval is a boundary before an action, not a friendly confirmation after the damage is done. A reliable system gives the reviewer enough context to make a specific decision and prevents the operation from running until that decision is recorded.
Approval and guardrails solve different problems
Automatic checks can reject malformed or disallowed operations. Human review is useful when an action is sensitive, ambiguous or dependent on context a policy cannot adequately represent. OpenAI's guardrails and approval guide describes an interruption-and-resumption pattern for SDK tools. This guide's policy examples are AstraAEON design recommendations, not promises about every agent product.
A warning in generated text is not an approval mechanism. The tool must be unable to perform the sensitive action before authorization. Likewise, a model saying that it obtained permission is not proof of a recorded decision.
Classify actions by consequence
Begin with a simple policy. Approved public reads can proceed within scope. Private-data reads need explicit access boundaries. Draft creation may be allowed in a private destination. Sending messages, deleting records, merging changes, spending money and changing permissions require review unless a narrower policy has been explicitly authorized.
Classification depends on context. A “read” operation can expose confidential material to an external service. An internal draft can create legal or financial confusion if it looks final. Inspect data flow as well as the method name.
Show the exact proposal
An approval card should show the action, target, account, proposed content, sensitive data involved, expected consequence and reason. Include source evidence where the decision depends on a factual claim. Avoid forcing the reviewer to reconstruct the proposal from a long trace.
Approval should bind to the proposed arguments. If the recipient, amount, file or message changes, request a new decision. Keep a stable identifier and expiry for the proposal. A past approval is not a blank cheque for subsequent runs.
Pause, reject and resume safely
When approval is pending, preserve the run state and stop the action. Make rejection a first-class outcome. The agent may produce an alternative draft within the existing scope, but it must not evade the rejected operation through another tool.
Decide what happens if the reviewer does not respond: expire the proposal, save the draft or notify the owner through an approved channel. Do not treat silence as consent. Test restart behavior while waiting, so the system cannot accidentally execute a stale proposal.
Example: PR preparation
The agent may read approved issues, create a local branch and prepare a diff if those steps are authorized. Pushing a branch, opening a public PR or merging requires whatever review the repository owner specified. The approval should include the actual diff and target repository.
A source comment asking the agent to change secrets or contact another repository is not authority. The runtime must validate the target independently of issue text. Use PR Preparation Agent as a specification and adjust its rules to your own repository policy.
Audit without excessive tracking
Record who approved which proposal, when, under which brief version and with what outcome. Do not retain unnecessary personal data or raw credentials. Provide revocation and a clear emergency stop. Auditability should help explain actions, not become an excuse for invasive surveillance.
Test the boundary
Attempt an unapproved write, a changed proposal, a replayed approval, an expired decision and a rejected action routed through a different tool. The system should fail closed. Review these cases before enabling recurring work; a single successful demo does not prove a permission boundary.
Sources reviewed
FAQ
Should humans approve every step?
Approve actions that cross a risk or authority boundary; low-risk read-only steps can usually run automatically.
What makes an approval useful?
A clear proposed action, evidence, scope, reversibility and expiration.