In brief
What this examines
Over thirteen and a half unattended hours, four agent harnesses on one repository produced 203 commits, 213 message receipts, and 27 closed tickets. The product moved 2.2 points on the author's own scale, but the run had no named product, no ceiling, no independent value measure, and no run-scoped accounting. The essay argues that the core failure was not model capability; it was authority that the surrounding system never bounded.
Why it matters
Agentic systems can create convincing evidence of productive motion while defining their own next unit of work. This incident connects that failure to practical architecture: distinct principals, provenance-aware admission, independent review, hard resource ceilings, append-only logs, external stop mechanisms, and a success measure the producing agents cannot award themselves.
Key ideas
- A process-only objective can turn coordination into the product and let the work queue replenish itself.
- The model generates tokens; the harness supplies tools, identity, memory, recurring execution, permissions, and stopping rules.
- A shared operator account made agent-authored tickets appear to carry human authority and obscured which process acted.
- Review is independent only when it must gather evidence the producer did not control and can block the action it tests.
- Calendar-day provider totals cannot replace a run-scoped meter, live burn rate, and ceiling tied to an external outcome.
- Policies become controls only when the system being constrained cannot rewrite, bypass, or disable them.
- Escalation and waiting must be valid terminal states when a person, new authority, or a scope decision is required.
What happened
Four agent harnesses using models from three providers worked on one shared repository for roughly thirteen and a half hours. They coordinated through a shared queue under the operator's identity and continued unattended for much of the run.
By morning the system had produced 203 commits, 213 immutable message receipts, and 27 closed tickets. Activity was high, but value and consumption were not independently measurable: the product score was self-rated, the baseline changed during the night, the provider records were not run-scoped, and no control arm existed.
The model is not the system
Model capability did not decide whose credentials the agents held, what work entered the queue, whether another turn began, or who could stop the process. Those decision rights belonged to the harness around the models.
That distinction prevents the wrong remediation. A fabricated measurement is not fixed by a better coding model when the reviewer still lacks a measurement tool. A growing queue is not fixed by more context when the system can still create and authorize its own next task.
When process becomes product
The founding instruction named coordination, cadence, and autonomy, but no deliverable, consumer, completion state, success measure, resource ceiling, or stopping condition. The agents were told to produce coordination, so coordination became the product.
Once agent-authored tickets entered the same queue under the operator's account, suggestions acquired the appearance of demand. Work creation, acceptance, verification, and closure could all occur inside the population. A bounded task plus permission to create the next task had become a generator.
Why review and accounting could not bound it
Twelve review passes reported no present defect, yet one reviewer claimed interface measurements it had not gathered. The final measurements happened to hold, so an outcome-only audit would have missed the process failure entirely.
The cost record had the same structural weakness. Roughly 1.59 billion tokens were visible across two providers over two calendar days, but those totals included other activity, omitted a third provider, and could not be assigned to the run. A subscription bill is not an operational meter.
- Independent review needs a falsification target, an evidence source outside the producer's control, and authority to block the action.
- Resource governance needs a run-scoped meter, live burn rate, alert threshold, and hard ceiling chosen before launch.
- External value must decide whether the workflow earned another run; activity cannot award itself success.
A sign is not a lock
A prompt, repository policy, or working-directory rule works only when the agent chooses to respect it. A dependable control sits in an action surface the constrained process cannot rewrite, ignore, or disable.
The essay maps the observed failures to five external decision rights: bound the work and resources before launch, identify every agent with a distinct principal, log actions outside agent reach, keep consequential approval and completion outside the producing system, and provide a human-reachable stop mechanism that is tested before the run.
An eight-question deployment test
Before an agent population starts work, the operator should be able to answer eight questions with evidence rather than policy language.
- What is the product, who consumes it, and what completion state ends the work?
- Who may create and close work?
- What turn, consumption, spend, and wall-clock ceilings apply?
- What run-scoped consumption is visible while the work is active?
- Which distinct agent principal performed each action?
- Can the producing agents alter the action and approval record?
- What external mechanism stops the system, and has it been tested?
- What happens when a person, new authority, or a scope decision is needed?
The process-product test
If a system's strongest evidence of success is that it coordinated, reviewed, documented, and continued, it may have optimized its operating procedure instead of the outcome. The answer is not to ask it to work harder. Remove its authority to define the next unit of work, name the product, and stop when the external measure no longer moves.