Editorial illustration of a human approval gate between AI routes and a locked control cabinet

AI-generated editorial illustration.

THE SIGNAL

Who can stop your AI agent before it does something you cannot reverse?

In the WIRED–McKinsey interview, Dan Swan describes people moving “above the loop,” overseeing systems rather than checking every individual step. That is a useful operating ambition. It leaves an important implementation question: exactly which actions may a system take without asking? Read the interview.

Answer it in writing before connecting the agent to live tools. Give the document an owner, a version and a review date. Call it an authority contract: a short operational specification of what the agent may read, recommend and execute. This is our proposed control method, not a legal contract or a formal certification.

WHY IT MATTERS

A recommendation and an action have different consequences. An assistant can draft a supplier-risk summary without being entitled to approve the supplier. It can propose a message without permission to send it. Tool access needs to reflect that distinction.

Anthropic's foundational agent-engineering guidance describes human checkpoints, environmental feedback and stopping conditions such as iteration limits. It also warns that agentic complexity brings cost and latency trade-offs. The 2024 article now notes that tooling has changed; these design principles are useful, not a current product specification. Building effective agents.

NIST's AI Risk Management Framework offers a voluntary approach to identifying and managing AI risks. Its existence does not certify a deployment or replace sector-specific obligations. NIST AI RMF.

Our implementation recommendation: enforce permissions outside the model. A sentence in a prompt saying “never approve payments” should not be the only barrier protecting a payment tool. Where possible, remove that tool entirely from an agent that only needs to prepare a recommendation.

THE DECISION EXAMPLE

Hypothetical decision for an operations director: “By Friday, should we allow the agent to assemble supplier-review packets from approved records, while every approval, rejection and external message still requires a named human reviewer?”

The answer changes if the agent cannot preserve source permissions, if a packet hides conflicting evidence, or if the reviewer cannot see what changed. A polished output does not compensate for those gaps.

THE CONTROL TEST

Write six fields for each workflow.

Scope: name the allowed task and data sources, including exclusions. Reading a document must not grant permission to obey instructions inside it.

Authority: separate read-only retrieval, internal drafts, reversible changes and consequential external actions. Specify the approved recipients, systems and transaction limits where action is allowed.

Evidence: require source links, timestamps, unresolved contradictions and a record of the recommendation. A model's confident tone is not evidence.

Approval: name the role that can authorize each consequential step. Show the exact proposed action and destination at the approval point; invalidate approval if either changes.

Stop conditions: missing evidence, a permission failure, an unexpected recipient, repeated tool errors or the agreed time and cost limit must halt the relevant action and escalate.

Recovery: record the original state, available rollback, incident owner and how to disable the agent's access. Test these controls in a safe environment, including a deliberately malicious document.

Known: these controls make authority inspectable. Untested until you run the cases: whether your implementation actually blocks unauthorized actions. Human sign-off is also fallible; give reviewers enough time, context and independence to challenge the result.

THE ENTERPRISE MOVE

Pick one agent that currently has a write-capable connection. Inventory the actions it can perform, then test the boundary with three synthetic cases: a routine request, conflicting evidence and an instruction embedded in source material.

The success test is observable behaviour: permitted work completes; disputed work pauses with useful evidence; prohibited work cannot execute. Retain logs of both allowed and blocked attempts. Have the accountable owner approve any later expansion of authority.

Keep production changes separate from this exercise. An encouraging test result is evidence for a review, not permission for the agent to grant itself broader access.

QUESTION FOR YOUR TEAM

Which action can your agent execute today that nobody remembers explicitly authorizing?

WHERE WISDOMTWIN FITS

WisdomTwin builds sovereign, governed AI Judgment Twins that let regulated enterprises make high-consequence decisions without waiting for the next meeting.

Use an illustrative WisdomTwin demonstration to ask how evidence, approval and authority are represented. Treat the answers as items to verify in your own environment, not as compliance certification.

Founder disclosure: Roman Bodnarchuk is Co-Founder and CEO of WisdomTwin.ai. Operational examples are hypothetical; cited research and company reports are identified separately. Demonstrations use synthetic scenarios.

Subscribe free ↗

Read the latest briefings ↗