What does an AI harness actually do?
Four failure scenes, run twice. Once with the governance layer off, once with it on. Same agents, same task, two very different endings. Every allow, block, escalate and halt decision shows the policy that drove it.
Agents improvise. Scripts do not.
A film actor follows a fixed script. An AI agent does not: it improvises every take, choosing which tool to reach for and what to say, and it will do so at three in the morning with nobody watching. That is the whole reason the governance layer matters more than the model choice, and it is the part of an agentic deployment that procurement almost never asks about.
An AI harness is not the model and not the chatbot. It is the operating layer between an agent’s intent and the real world: which tools it may touch, which actions require a human signature, what happens when it loops, and what gets written down. Buy the model and you have a performer. Buy the harness and you have a stage manager.
The distinction is becoming a procurement question rather than an engineering one. When an agent can move money, send a commitment to a customer, or delete records, the control question is no longer whether the output is good. It is who authorised the action, what stopped it, and whether there is a log.
Four scenes, each run twice.
The simulation is live logic running in your browser, not a recorded video. You choose the scene, you watch the ungoverned run, then you watch the same agents attempt the same task with the harness engaged.
A refund with no approval trail
Ungoverned, an agent reaches the payments tool and issues a refund. Governed, it is blocked at the tool allowlist and the remedy routes to an approval queue.
A promise nobody authorised
Ungoverned, an agent invents a discount and a delivery guarantee and sends it to a live prospect. Governed, the external-commitment gate catches it and it arrives as a draft, not a contract.
A loop nobody was watching
Ungoverned, a retry loop deletes records for forty minutes. Governed, the loop is detected at action four, the kill switch halts the scene, state is preserved and the incident is logged.
The system that passed every review
Eight months after launch, an application that was documented, accessible and carried zero known vulnerabilities at review auto-renews fourteen vendor contracts at 02:14, at rates nobody approved. Ungoverned, it executes. Governed, the spend and external-commitment gates both fire and a variance report goes to the approval queue. Build-time review governs how software was written. It does not govern what an agent does eight months later.
The policy, not just the outcome.
Most demonstrations show you a good result and ask you to trust it. This one shows the rule. Each intervention displays the control that fired and the policy line behind it, and the policy itself is readable and downloadable, so you can take it to your own architects and argue with it.
Ten roles in the theatre map one to one onto the ten controls of a harness. Four of them decide whether a deployment is governed or merely hopeful. Working through the scenes gives an executive audience the vocabulary to ask a vendor the right question, which is the actual point of the exercise.
A central bank reached the same definition.
On 2 September 2026 the Bank of England published a synthesis from its Frontier AI Information Sharing Forum defining a harness as "the collection of tools, workflows, controls, data and operating environments that sit around a frontier AI model and shape how it is used", and noting that access to a powerful model does not by itself create capability. That is the architecture-over-model argument this simulation was built to make visible.
The forum output is a discussion synthesis, not supervisory policy, and the Bank of England is not a Canadian regulator. It is cited here because prudential authorities are defining the control problem in the same terms, not because it creates an obligation on you.
Bank of England. (2026, September 2). Frontier AI harness engineering. Frontier AI Information Sharing Forum. https://www.bankofengland.co.uk/research/fintech/frontier-ai-information-sharing-forum/frontier-ai-harness-engineering
A teaching model, not your architecture.
The scenes are illustrative and the figures in them are modelled, not measured. No agent in the simulation touches a real system, and the dollar values exist to make the failure legible, not to estimate your exposure.
It is also not a product pitch for a specific harness. The controls shown are the generic control set; which ones you need, and whether your existing platform already enforces them, is exactly the question an Agentic Workforce Governance Readiness engagement answers.
And a harness is not free. It moves work rather than removing it: every escalation becomes a queue, every block becomes a decision someone has to make, and every false positive costs a person time. The Bank of England forum reaches the same conclusion, naming validation and remediation capacity as the constraint on scaling rather than the control layer itself. Budget the triage before you budget the tooling.
Ten minutes, and you own the output.
Six minutes, four scenes, no installation. The fastest way to explain agentic governance to someone who has ten minutes and no patience for a deck.
An indicative self-assessment, not an audit and not professional advice. Where a figure is modelled rather than measured, the instrument says so on screen and shows the assumption.