Most failures I encounter in agent-assisted work are not failures of intelligence. They are failures of control.
The artifact is downstream
A model can write good code against the wrong repository, summarize a stale document accurately, or produce a convincing completion note for a change that was never verified. The artifact may be competent while the operation is still wrong.
That distinction matters because prompts optimize the immediate response. A harness governs the conditions under which the response becomes an action, a claim, or a release.
The structure is cybernetic
Evidence Harness follows a feedback loop: establish the desired state, sense the actual state, compare the two, permit only a bounded action, observe the result, and feed demonstrated failures back into the rules. The receipt closes the loop by making the comparison inspectable.
Route → authority → sense → bound → verify → receipt → improve
The most important behavior is sometimes to do nothing. If authority is missing, state is contradictory, or permission is absent, the truthful output is blocked or unknown—not a synthetic green.
Human authority remains part of the system
A useful harness distinguishes observation from mutation and local work from external consequence. Approval to draft is not approval to deploy. Approval to prepare a release is not approval to publish it. This is not ceremony; it is how an agent remains legible to the people responsible for the outcome.
Receipts change the meaning of “done”
Completion should name the evidence that supports it: the test that passed, the live endpoint that resolved, the exact commit that shipped, or the limitation that remains. A receipt does not manufacture proof. It records whether the proof exists.
This makes the operating system improvable. When a real failure repeats, you can ask whether the rule was missing, whether routing failed to load it, or whether enforcement ignored it. The response is a narrow control and, when possible, a behavioral test—not a larger prompt full of accumulated anxieties.
Why release it
Evidence Harness is a sanitized public extraction of the control layer I use in my own stack. It contains no private prompts, production identities, client material, or operational traces. Its first proof is deliberately synthetic: a project mismatch and missing authorization block a deployment before any command runs.
The release is a preview, not a claim of universal safety or third-party outcomes. Its value is simpler: the operating pattern is now inspectable, installable, testable, and open to improvement.