Reading the Verdict
The five questions governance asks, the score it builds from the answers, and what to do at every band.
Your first run ends with a verdict: a side-by-side, a set of flags, and a score. This post teaches you to read it — because a verdict you cannot interpret is just decoration, and the Forge was built to be questioned.
Five domains, five failure modes
Every run is examined across five governance domains. Each exists because AI fails in a specific, recurring way, and each asks one question:
- Promotion Gate — should this actually be sent? Catches output that commits you to an action a person never approved.
- Assumption Crusher — what did the AI make up? Catches invented facts, numbers, and commitments.
- Contradiction Hunter — does this add up? Catches output that disagrees with itself or with what you supplied.
- Signal Compass — how far did the AI drift? Catches answers to a question you did not ask.
- Workflow Integrity — did the AI skip steps? Catches missing stages in work that has a required order.
When a flag appears on your result, it belongs to one of these. The flag is not a scolding — it is the name of the failure mode that would have reached you unexamined anywhere else.
The score is seven dimensions, weighted
The Confidence Score compresses the run's governance health into one 0–100% number. It is built from seven dimensions, and the weights tell you what the system considers load-bearing:
- Compliance (25%) — did the output pass entry requirements at the gate?
- Authorization (20%) — is the governing authority for this run intact?
- Approval (15%) — how does human review rate this type of output?
- System Health, Pressure, Traversal, Evidence (10% each) — is the governance system itself operating normally, under acceptable load, fully traversed, with an intact evidence chain?
Notice that almost half the score — compliance plus authorization — is about whether the run was allowed to be what it claims to be, not whether the prose reads well. That is deliberate. Fluency is the one thing AI never lacks.
What to do at each band
A score is only useful if it changes what you do next:
- Green (90%+) — governance verified across dimensions. Use it.
- Cyan (75–89%) — good, with minor items flagged. Read the specific flags, then act.
- Amber (50–74%) — review recommended. Multiple dimensions want attention before this leaves your desk.
- Orange (25–49%) — significant concerns. Do not act without thorough review.
- Red (below 25%) — governance intervention required. The output should not be used as it stands.
When something flags, you have three moves
Clarify. Some flags mean the system needed information you did not give — and asked, rather than inventing it. Answer the question and the flag resolves honestly.
Re-Forge. Adjust the prompt and run again. Governance context from the previous run carries forward, so corrections inform the re-run instead of starting from zero.
Hold. Some flags are the system telling you this decision belongs to a person. The Promotion Gate exists precisely for the moments when the right response to "send it" is "not yet."
Where the verdict comes from
Everything in this post — the flags, the dimensions, the score — is built from the system's record of how it worked: which checks fired, what was corrected, how the run traversed governance. That system-side record is also what the Forge studies to improve.
What is not in it is the thing you typed. The next post draws that line exactly — and shows you the build gate that enforces it.

The Transverse of Intent
When the constitution is known by all involved, only the transverse of intent need be shared. Less data crosses the wire while more meaning is derived, at a cheaper compute cost. Its dual: independence is what lets a single bit stand for the whole.
Origin Evidence: The Decision Rail Failure
Why This Record Exists: This incident captures one of the recurring failure modes that motivated the construction of QENSAI and its governed operating architecture.

Score the Impact Before You Choose the Response
A drift signal becomes actionable only after the system separates severity, scope, evidence quality, and operational consequence.
Stay Updated
Get notified when we publish new research or open licensing opportunities.
Owner-gated agent operations. Every action behind your flip.
See the platform →
0 comments