James F. GibbonsEnterprise Search & Applied AI

Agent operations · August 2026

Govern the harness, not the artifact

An AI tool can write good code against the wrong project—or say it finished before anyone checked the page. The working process around the tool matters as much as the output.

Most failures I encounter in agent-assisted work are not failures of intelligence. They are failures of control.

The artifact is downstream

A model can write good code against the wrong repository, summarize a stale document accurately, or produce a convincing completion note for a change that was never verified. The artifact may be competent while the operation is still wrong.

That distinction matters because prompts optimize the immediate response. A harness governs the conditions under which the response becomes an action, a claim, or a release.

The structure is cybernetic

Evidence Harness follows a feedback loop: establish the desired state, sense the actual state, compare the two, permit only a bounded action, observe the result, and feed demonstrated failures back into the rules. The receipt closes the loop by making the comparison inspectable.

Route → authority → sense → bound → verify → receipt → improve

The most important behavior is sometimes to do nothing. If authority is missing, state is contradictory, or permission is absent, the truthful output is blocked or unknown—not a synthetic green.

Human authority remains part of the system

A useful harness distinguishes observation from mutation and local work from external consequence. Approval to draft is not approval to deploy. Approval to prepare a release is not approval to publish it. This is not ceremony; it is how an agent remains legible to the people responsible for the outcome.

Receipts change the meaning of “done”

Completion should name the evidence that supports it: the test that passed, the live endpoint that resolved, the exact commit that shipped, or the limitation that remains. A receipt does not manufacture proof. It records whether the proof exists.

This makes the operating system improvable. When a real failure repeats, you can ask whether the rule was missing, whether routing failed to load it, or whether enforcement ignored it. The response is a narrow control and, when possible, a behavioral test—not a larger prompt full of accumulated anxieties.

Why release it

Evidence Harness is a sanitized public extraction of the control layer I use in my own stack. It contains no private prompts, production identities, client material, or operational traces. Its first proof is deliberately synthetic: a project mismatch and missing authorization block a deployment before any command runs.

The release is a preview, not a claim of universal safety or third-party outcomes. Its value is simpler: the operating pattern is now inspectable, installable, testable, and open to improvement.