The governed intelligence playbook
The plain version first: these are the rules I build into AI systems so any operation can use them without losing control of its own decisions. Below - how the decision loop gets its intelligence layer, and the primitives that keep it accountable, starting with the one that prevents a breach by architecture, not by policy.
The Intelligence Stack
Six layers carry the decision loop from raw signal to bounded purpose. A single governance layer runs through all of them - not bolted on after the fact, built into the structure itself.
Select any layer, or the governance bar, for detail. Escape or the close button collapses it again.
Stack layers, bottom to top: Learn, Orchestrate, Decide, Interpret, Sense, Purpose
The loop runs continuously - Learn feeds its corrections straight back into the next cycle of Sense.
Layer detail
We already run this loop today, on humans and Power BI. The gold bar is what is missing. That is what I am proposing to build.
The Meta Rule of Two
No single agent session may hold all three: untrusted input, private data access, and an external outbound channel. Any two is contained. All three is a breach waiting to happen.
Illustrative - not a scored claim. This is a working simulator of the architecture rule, not a live system status feed.
Grant capabilities to the agent session
- Untrusted input
- Private data access
- External outbound channel
Resulting architecture state
Session diagram
Status
Idle. No capabilities granted - nothing to contain.
Evals, logs, rollback, and the human review queue are detection and recovery - they fire after something has gone wrong. The Rule of Two is prevention by architecture. It is the fifth primitive.
Their four primitives tell you a breach happened and help you undo it. The Rule of Two makes the breach structurally impossible. You want both.
The fiduciary wedge
Every decision below has a side. The wedge is the permanent gap between them - it is what keeps a routine reorder and a customer-facing call from ever being routed the same way.
The two peaks
HIGH-SIGMA -> HUMANS
Ambiguous, high-stakes, value-laden decisions stay with named humans.
LOW-SIGMA -> AGENTS
Routine, reversible, well-bounded decisions route to agents under audit.
A named human always stands behind the consequential decision.
It is called a wedge because the gap between what AI can do and what it can be held accountable for is permanent. It does not close as models improve. That is why governance is architecture, not a feature.
Route a decision
Drag a decision chip onto a peak, or select a chip and choose where it routes.
The Exception Gate
A wide stream of decisions narrows through a single gate. Most continue on their own. A named minority are pulled aside before anything executes.
Illustrative - not a scored claim. This is a working simulator of the escalation rule, not a live system status feed.
The rule
Any decision touching money, legal text, or customers-of-record automatically escalates to a named human.
Decision flow
Static view - motion is reduced or canvas is unavailable. Same split, same rule.
Escalation controls
Hover or focus the gate mark to view what triggers escalation.
At this threshold, 5% escalate to a Named Human Fiduciary and 95% continue as Autonomous Execution (Low-Sigma).
illustrative - not a scored claim
Automation handles the expected. Humans handle the ambiguous.
Bolted-on vs designed-in
The same organization, two ways of bringing in intelligence. One is stapled onto the org chart that already exists. The other sits exactly where the real work already crosses.
Drag the divider, or use the arrow keys, to compare. Or jump straight to either side below.
Jump to a view
Bolted-on and designed-in, side by side
Current view
Both views side by side: automation stapled onto an unchanged org chart on the left, synthetic nodes woven into the flow on the right.
Bolted-On: automating the wrong things. Leaving the deepest friction untouched.
Designed-In: synthetic nodes at the structural intersections.
You cannot achieve an intelligence-led transformation without an honest accounting of the legacy architecture you are building on.
Bolted-on automates the wrong things and leaves the deepest friction untouched. My team builds at the intersections because we are the intersections.
The surgical undo.
A version timeline for one agent - prompt, model and policy versioned like software, so a bad change can be undone at the scope of one decision, not the whole stack.
Illustrative - not a scored claim. The version timeline, the changelog entries and the violation state below are a working simulation of a rollback flow, not a live deployment log.
Version timeline
Changelog
Select a version above to see its changelog entry.
Rollback status
Stack running at v1.3, stable. Press Trigger the violation to simulate a policy violation and its rollback.
Revert an agent to last week's prompt, last month's model, or last quarter's policy version without taking the stack down. Kill-switch thresholds and soft-delete windows on every destructive endpoint.
Agent versions are software versions: traceable, diffable, recoverable.
The Organizational Equalizer
Seven structural dimensions, one board. Move any fader from Legacy toward AI-Native and watch which two dimensions the whole system is actually waiting on.
Illustrative starting positions - not a scored claim. Julia's honest scoring lands here after her workshop pass.
illustrative - not a scored claim
Look strictly at your two lowest dimensions. Every other initiative is gated by them.
Total 33 out of 70. Foundational Work Needed.
I am not asking for a role. I measured the organization. These two dimensions gate every other initiative, and the measurement points at the fix.
The readiness matrix
Score an agent class on the four detection-and-recovery primitives, one to five, from the legacy firm to the ExO 3.0 firm (the exponential-organization maturity model at full score). Then check the one primitive that allows no partial credit.
Score this primitive from 1 to 5
Evals
Not yet scoredScore 1 - The Legacy Firm
Manual QA before launch, blind in production.
Score 5 - The ExO 3.0 Firm
Continuous synthetic evaluation; auto-halt on drift.
Logs
Not yet scoredScore 1 - The Legacy Firm
Scattered chat threads and isolated app logs.
Score 5 - The ExO 3.0 Firm
Immutable correlation IDs spanning the entire loop.
Rollback
Not yet scoredScore 1 - The Legacy Firm
Shut down the entire server to fix the agent.
Score 5 - The ExO 3.0 Firm
Surgical prompt and policy reversion with zero downtime.
Human Queue
Not yet scoredScore 1 - The Legacy Firm
A human in every loop, crushing speed.
Score 5 - The ExO 3.0 Firm
Humans above the loop, SLA-driven exception management.
Rule of Two compliance
No agent session holds untrusted input, private data access, and an outbound channel at once.
Pass or fail only. No partial credit.
Rule of Two compliance status, pass or fail
Your scores, stored locally in your browser - illustrative, not a published claim.
Do not deploy a new agent class. Failing: Evals, Logs, Rollback, Human Queue, Rule of Two.
Do not deploy a new agent class until you score a 3 across all four rows - and the Rule of Two is pass/fail.
Framework vocabulary on this page draws on: openexo.com/book, openexo.com/resource-hub
Operate the systems this governs. See how autonomy is earned.