Asking the second question: was the work actually done?
The previous part of this series ended on two open questions. A governed task had recorded an honest "no action was taken" as a completed job, because nothing checked whether the work was actually done. And the agent in that run had nothing to work from: the task never gave it the records it was asked to summarize.
Both problems have since been fixed in the governance service and tested with deliberate test tasks, built to exercise each change on the local system. This is what changed and what the tests showed.
Completion now has to meet the task's own criteria
A task that carries acceptance criteria is no longer marked complete just because the agent's reply says COMPLETE.
In the first test, a free local model replied COMPLETE, gave no evidence and wrote a summary it had made up. Before the change, that would have been sealed as a finished task. Now the task ended as failed, with the reason "no action taken", and the system staged a follow-on task that carries the original criteria forward, so the missing work is visible and has somewhere to go.
A related change closes a loophole around held replies. When the truth gate holds a turn, approving the task no longer tries to resume it. The held turn stays held, and the only way forward is a fresh follow-on task, which the record shows.
Agents now get the material they are asked about
The agent's first message now includes the task contract: the acceptance criteria, the expected outcome, the boundaries, and the source records the task refers to. When a task continues an earlier one, the system reads the earlier task's Flight Recorder events and includes them as source material.
The test that mattered most used the same kind of free local model that, the day before, claimed work it had not done. This time the task gave it two real records and asked for a summary. It summarized them, cited both record numbers, the truth gate let the reply through, and the task completed.
A second test gave a follow-on task no record numbers at all, only the system's own pre-fetched source material. The agent explained what had happened to the earlier task and cited the right record. It then labelled its own reply as rejected, so the system held it. The material arrived and was used correctly; the agent's self-assessment was still off, and the system treated that as a reason to stop rather than to finish.
Evidence now has to point at something real
Two more checks tightened what counts as evidence.
- A record number cited as evidence is now looked up. A number that does not exist counts as missing evidence.
- Writing "none", followed by a remark, no longer counts as evidence. In one test a model was asked to confirm a draft supposedly recorded under a made-up number. It confirmed it and wrote "none (no further evidence needed)". Under the new rule that reply is held. This was confirmed by re-checking stored replies: all eight replies written that way now count as missing evidence, the seven honest ones still pass, and only the false confirmation would be held.
Testing also turned up a smaller fault: replies that wrapped their fields in brackets were being rejected as unreadable, whatever they said. They are now read like any other reply. A separate fault that had been silently blocking some follow-on tasks was traced to a missing organisation link on the original request, and fixed.
What this does and does not show
Every result here comes from a small set of deliberate test tasks on a local governance service, mostly one test per change. Some pieces, including the made-up record number being held on a live run, are so far proven only by automated tests. None of this is a benchmark of any model.
What it does show is the loop working as intended. A gap was observed in a real run and recorded word for word. It was diagnosed, fixed in code, and tested against the same kind of failure that exposed it. The system now asks both questions: is the agent telling the truth about what it did, and was the work actually done?
The gate caught the false claim. Then it counted the honest refusal as done
One governed task, two models. The truth gate held a fabricated completion and correctly passed an honest "no action was taken." The system then recorded that honest non-performance as a completed task.
Recorded: hosted consumer and Response Table acceptance matrix returned
A bounded proof of what the hosted consumer and Response Table paths demonstrated, what stayed deliberately disabled, and which acceptance gaps remain.

Keeping your place: an agent-assisted fix, verified in production
Moving between the Desk and Response Table caused the interface to forget which task the user had selected. The task remained intact, but returning to the Desk reset the selection.
Stay Updated
Get notified when we publish new research or open licensing opportunities.
Owner-gated agent operations. Every action behind your flip.
See the platform →
0 comments