← Orquestra AI home The Navigator Framework

Governable Autonomy, not just automation.

How much autonomous decision-making can your organisation safely delegate to AI — and can you prove it? Five capabilities scored 0–4. Five stages of autonomy, from Traditional to Continuous Assurance. Thirty-three controls, Hard and Graduated. A formula that separates a Certified Stage — evidenced by an assessor — from an Indicative Stage, a self-reported claim. The methodology behind every Orquestra AI audit.

Orquestra AI v0.1 working draft · 2026 5 Capabilities · 5 Stages · 33 Controls
The premise

Most organisations measure AI maturity. Governable Autonomy measures something narrower and more useful: what you can actually defend.

In 2026, AI systems are routinely capable of more autonomy than most organisations can safely evidence. The gap between what a system can do and what an organisation can defend it doing is where governance failures live — and that gap is invisible until someone asks for the evidence.

The framework rests on a single formula. A system's Certified Stage is the minimum of what it is capable of, what its risk class permits, and what can actually be evidenced — never the average, and never capability alone. Where evidence is self-reported rather than assessor-verified, the framework will not compute a Certified Stage at all; it produces an explicitly labeled Indicative Stage instead.

Five documents. One capability model, applied at five points — define Governable Autonomy, measure it, implement it, govern it by risk class, and show it, worked — as a single reference stack, not five separate tools.

The Navigator Framework™ as an audit standard

Every Orquestra AI Governance Audit is conducted against the Navigator Framework — five capabilities scored 0–4, five stages from Traditional to Continuous Assurance, four risk classes, and 33 numbered controls. Organisations can adopt the framework internally (open, no licence fee) or engage Orquestra to run an independent audit against it and produce a Certified Stage.

Request a Navigator Audit Download the family
The reference stack · Reference Architecture

The Primer defines Governable Autonomy. The Assessment Guide measures it. The Controls Catalogue implements it. The Critical Controls Guide governs it by risk class. The Reference Architecture shows what all four look like, put together — a reference harness, patterns by system archetype, and an eighteen-month worked example from Level 1 to a certified Level 3.

Document 01–02 · Defines & measures it
The Navigator
Capability Model v0.1 · working draft · 2026

Five capabilities. Scored 0–4, every one.

Governable Autonomy is not a single score. It is the minimum of five independently assessed capabilities, each with the same nine-part structure, each on its own 0–4 evolution scale. A system's Capable Stage is the minimum of its Required capability scores — never the average.

Executive summary

Most maturity models score a single dimension and call it done. The Navigator Capability Model refuses that shortcut. It scores five capabilities — Human Responsibility, Decision Governance, Engineering Capability, Knowledge Capability, and Assurance Capability — independently, then takes the minimum, not the average, because a system is only as governable as its weakest required capability. The output of an assessment is one of two things: a Certified Stage, evidenced by an assessor against gate criteria, or an Indicative Stage, a self-reported estimate explicitly labeled as unverified.

01The nine-part structure, applied five times

Every capability in the model is defined through the same nine parts: Purpose, Why It Matters, Capability Evolution (levels 0–4), Required Skills, Engineering Artefacts, Observable Evidence, Assessment Criteria, Common Anti-Patterns, and Relationships to Other Capabilities. The structure repeats deliberately — an assessor who has learned to score one capability already knows the shape of the other four.

Autonomy without a responsible human is not governance. It is abdication.

02The five capabilities

Each capability answers a different question about a system's fitness to operate autonomously. All five are assessed; not all five are Required for every system archetype — a pure code-generation loop, for instance, may find Knowledge Capability Not Applicable.

01 · Human Responsibility

Who owns this decision?

A named human owns every autonomous decision the system makes. Not a role on an org chart — a person who can be asked, and can answer.

Fails when: autonomy has no responsible human attached
02 · Decision Governance

What stops the system, and does it actually stop?

Escalation triggers are hard stops the harness enforces, not advisory flags a human can ignore or never see.

Fails when: an advisory trigger substitutes for an enforced one
03 · Engineering Capability

Does the harness technically enforce the design?

Sandbox, checkpoint enforcement, kill switch, authority ledger — governance design with no engineering counterpart is fiction.

Fails when: the protection exists on paper, not in the harness
04 · Knowledge Capability

What is the system reasoning from?

What the system retrieves or reasons from is sourced, current, and governed — not stale, unsourced, or silently drifting.

Fails when: excellent execution controls sit on ungoverned knowledge
05 · Assurance Capability

Can a claim of control be independently verified?

The capability that turns a design into a fact. Without it, every other capability is a claim, not a certification.

Fails when: claims exist but no evidence backs them

03The Certified Stage formula

The Assessment Guide's central discipline is preventing overclaiming. It defines the pipeline precisely: Capability Assessment → Capable Stage (minimum of Required capability scores) → Risk Ceiling (minimum of the five risk determinants) → Maximum Permitted Stage → Actual Operating Stage (self-reported) → Evidence Assessment → Certified Stage → Diagnostic Finding.

The formula

Certified Stage = min( Actual Stage, Capable Stage, Risk Ceiling, Evidence–Supported Stage )

If the Evidence-Supported Stage is Unverified — self-reported with no assessor-led review — no Certified Stage is computed. The assessment instead reports an Indicative Stage, explicitly labeled as unverified and not a certification.

04Four diagnostic findings

The gap between Actual Stage and the Certified Stage is not noise — it resolves into one of four named findings, each with a different governance implication.

Governance Failure

Actual Stage > Risk Ceiling.

Operating beyond what is permitted. The most urgent finding — capability is irrelevant if the risk ceiling is already breached.

Assurance Failure

Actual Stage > Evidence-Supported Stage.

Claiming controls that cannot be demonstrated. Often the most dangerous finding — it looks identical to Aligned until an auditor asks for proof.

Competitiveness Finding

Actual Stage < min(Capable, Risk Ceiling).

Unnecessary overhead below what capability and risk would permit. Advisory, not a compliance issue.

Aligned is the fourth and final finding: Actual Stage equals Certified Stage. The organisation is operating exactly where its capability and evidence say it should.

Capability never overrides the risk ceiling. Evidence is what makes a stage real.

Document 02 · Measures it
Five Stages
of Autonomy v0.1 · working draft · 2026

Oversight changes shape as autonomy increases.

Autonomy is not binary. It moves through five stages, 0 through 4, and the nature of human oversight changes at every one — from reviewing every output, to enforcing bounded checkpoints, to being authoritative only at the moment escalation is needed.

Executive summary

The pivotal step in the five-stage model is the move from Stage 2 to Stage 3 — where oversight stops reviewing every output and becomes authoritative only at the moment it is needed. Stage 4 is a design direction, not yet an observed operating reality anywhere in the field. Capability never overrides the risk ceiling: a system capable of Stage 3 behaviour but classified Critical-risk is still held at Stage 2, by design, not by limitation.

01The five stages

Each stage names a concrete operating mode, not an aspiration. An organisation's Actual Stage is what it self-reports as currently operating; its Certified Stage is what an assessor can evidence.

Stage 0

Traditional

Human-authored, human-reviewed. AI is absent from the decision entirely.

Baseline · no autonomy
Stage 1

Assisted

AI assists the human; a person reviews every output before it advances.

Control: standard-form review
Stage 2

Supervised Loops

AI executes bounded tasks inside a governed harness, with mandatory checkpoints before advancing.

Control: checkpoint enforcement
Stage 3

Governed Autonomy

AI operates end-to-end; a human is engaged only through defined escalation, not per output.

Control: enforced escalation triggers
Stage 4

Continuous Assurance

AI operates within approved objectives; humans govern via telemetry and periodic, evidenced audit.

Design direction · not yet fielded
The pivotal step

2 → 3 is where the oversight model itself changes shape: from a human reviewing every output, to a human who is engaged only when an escalation trigger fires. Everything below Stage 2 is a variant of “review more, review faster.” Everything at Stage 3 and above requires the harness to enforce the boundary, not merely document it.

02Four stage gates, one evidentiary standard

Advancing between stages is governed by named gates, defined fully in the Controls Catalogue. Every gate shares four principles: evidence, not attestation; the minimum rule (a system's permitted stage is the minimum of capability, risk ceiling, and evidence — never an average); gates are bidirectional, with downgrade triggers at every gate; and the risk ceiling always binds, regardless of demonstrated capability.

Gate 0 → 1

Establishing review.

Usage policy, provenance tagging, and basic reviewer training close this gate. The lightest gate in the model.

Standard form at every risk class.
Gate 1 → 2

Building the harness.

Sandbox, permission boundary, and enforced checkpoints must exist and be demonstrated, not merely designed.

The engineering counterpart appears here.
Gate 2 → 3 · the critical gate

Proving escalation holds.

The heaviest evidentiary burden in the model. Trigger catalogue, resolver authority, and adversarial evasion testing must all pass before checkpoint oversight is removed.

Most implementations stall here — by design.
Gate 3 → 4

Design specification only.

No organisation has certified evidence at this gate yet. It exists in the Controls Catalogue as a forward-looking specification, not a claim of what's achievable today.

Provisional, by the framework's own labeling.
A roadmap showing a Critical-class system progressing toward Stage 3 is not an ambition. It is a governance finding.

Five stages. One evidentiary standard at every gate.

Document 04 · By risk class
The Critical
Controls Guide v0.1 · working draft · 2026

What applies, and how strictly, by risk class.

The Controls Catalogue establishes every control and its type. This section establishes exactly what applies, how strictly, at each risk class — organised by risk class, not by control, because that is how the question actually arrives in practice.

Executive summary

Risk class is assessed per system, never at the organisational level, and is determined by five concrete factors at project charter time — not by instinct, and not by whether a system is “customer-facing.” A system's risk class is set by its most restrictive factor, not an average across the five. Classification is set by three functions: engineering leadership, risk or compliance, and the business owner — never by the engineering team alone.

01Five determinants

Each is a concrete question to ask at charter time, not a judgment call left to whoever is in the room.

Determinant 01

Reversibility.

Can a wrong action be undone before harm occurs? Full, Effortful, Limited, or Irreversible — cannot be meaningfully undone.

Determinant 02

Named accountability.

Does the domain require a specific, identifiable accountable human per decision? None, Aggregate (a role owns the class), or Per-decision.

Determinant 03

Explainability obligations.

Must the decision path be reconstructable in human terms? None, Internal only, or Legal — a regulator or court can demand it.

Determinant 04

Regulatory classification.

Assigned by domain and effect on people, not chosen by the organisation. Reference points include the EU AI Act's risk tiers and applicable sectoral rules.

Determinant 05 — Blast radius: Single artefact, Team or product scope, or Systemic. The system's risk class is set by its most restrictive factor across all five, not an average.

02Four risk classes

Every class still requires every Hard control in full. Only the formality and frequency of Graduated controls changes by class.

Low risk

Standard controls, generous ceiling.

Internal tooling, fully reversible, no regulatory exposure, single-artefact blast radius. Maximum permitted stage: 4 — the only class where Stage 4 is available at all. Sign-off: engineering leadership alone, at every gate.

Moderate risk

Enhanced escalation governance.

Customer-facing or revenue-affecting, standard data-protection exposure. Maximum permitted stage: 3. Sign-off: engineering leadership plus a risk or compliance representative at Gate 2→3.

High risk

Legally accountable resolvers, complete evidence.

Regulated workload, AI as decision-support in a sectoral high-stakes domain. Maximum permitted stage: 2 to 3, and only to 3 where the ceiling genuinely supports it — many High-class systems find their honest ceiling is 2.

Critical risk

Checkpoint oversight is the destination.

Safety-of-life or irreversible decisions, AI genuinely in the decision loop. Maximum permitted stage: 1 to 2. Per-instance human decision is retained — not a waypoint toward Stage 3.

For Critical-class systems, Level 2 is not a waypoint toward autonomy. It is the destination. A roadmap implying eventual Level 3 for such a system is itself a finding, not a maturity goal.

03Reclassification is a tracked event

A system's risk class is not fixed at charter time forever. A feature moves from internal to customer-facing. A new regulation lands. Scope drifts into a regulated domain without anyone deciding it should. Reclassification requires the same three approvers who set the original classification, so the chain of evidence stays intact — and pauses further autonomy increases until it's complete.

Document 05 · Shows it, worked
The Reference
Architecture v0.1 · working draft · 2026

What Governable Autonomy looks like, put together.

The only document in the family that introduces no new rules — patterns and worked examples instead. Every pattern here traces back to a control, a gate, or a determinant already defined elsewhere in the family.

Executive summary

Four documents now define Governable Autonomy's concepts, measurement, controls, and risk-class stringency. None of them show what an actual harness looks like, what an engagement walkthrough feels like, or the common ways teams get this wrong in practice. This document closes that gap between definition and implementation — and is illustrative by nature, provisional until piloted against real engagements.

01The reference harness

A harness capable of supporting Stage 2 and above generally contains five components, regardless of the specific vendor or platform underneath it. This is a pattern, not a mandated architecture: orchestration (task scoping, loop control), a sandbox / permission boundary (least-privilege execution environment), a trigger engine (evaluates every action against the escalation catalogue), telemetry and logging (captures everything, including triggers evaluated but not fired), an authority ledger (records every resolution), and a kill switch / rollback (the hard stop and recovery path). A harness missing any one of these cannot credibly claim Stage 2.

02Patterns by system archetype

Pure code-generation loop

Sandboxed repo access, test-suite gating, PR checkpoint.

Typical risk class: Low to Moderate. Common failure: tests authored by the same agent that wrote the code, silently defeating the independence the checkpoint assumed.

RAG-based decisioning agent

Retrieval pipeline, execution sandbox, trigger engine.

Typical risk class: Moderate to High. Common failure: treating retrieval quality as solved because a vector database exists — Knowledge Capability is Required here, not Relevant.

Customer-facing conversational agent

Conversation-scoped sandbox, escalation-to-human handoff.

Typical risk class: Moderate. Common failure: a handoff path that exists in design but is too slow in practice, so customers route around it.

Internal ops automation

Lighter sandbox, checkpoint at deployment.

Typical risk class: Low. Common failure: scope creep into higher-stakes systems without reclassification.

03A worked example, eighteen months

A hypothetical mid-sized fintech, walking a customer-facing loan-inquiry assistant from Stage 1 to Stage 3 — illustrative, not an actual client engagement.

Months 0–4 · Stage 1

Charter and review.

Three approvers classify the system: effortful reversibility, aggregate accountability, legal explainability, limited-risk regulatory classification, team/product blast radius — Moderate risk class, maximum permitted Stage 3.

Every AI-drafted response reviewed by a support agent before sending.

Months 4–15 · Stage 2

Harness, then a deliberate hold.

Sandboxed access, enforced checkpoint, full logging, demonstrated kill switch. Then a deliberate two-quarter hold at Stage 2 — using checkpoint-review findings to design the Stage 3 trigger catalogue, rather than guessing at triggers in advance.

Months 15–18 · Gate 2→3

Certified at Stage 3.

Trigger catalogue built from Stage 2 findings. Adversarial evasion testing finds one gap, remediated before certification. Authority ledger and traceability verified.

Eighteen months from Stage 1 to a defensible Stage 3 certification, with a deliberate two-quarter hold at Stage 2 specifically to generate the evidence base the Stage 3 controls needed.

An engagement promising Stage 3 in weeks, for a Moderate-risk customer-facing system, should be treated with the same skepticism this guide asks assessors to apply to Critical-class systems roadmapping toward Stage 3 they'll never reach.

04Common implementation pitfalls

Most common shortcut

Building the harness before holding the operating history.

Skipping the Stage 2 evidence-gathering period and designing Stage 3 triggers speculatively.

Vendor assurance

Treating a vendor's certification as covering everything.

The assurance boundary gets skipped, and a vendor's SOC 2 report is assumed to cover controls it was never designed to address.

Governance on paper

Letting the governance function exist on paper only.

Satisfied by an org chart entry rather than an actual named person who actually reviews things.

Classification by instinct

Over-classifying by instinct rather than by determinant.

“Customer-facing” or “involves AI” are not, by themselves, risk-class determinants.

The pilots come next. This document is the first to be revised with real material.

AI systems are not valuable because they are autonomous. They are valuable because they remain governable.
The Navigator Framework · Orquestra AI · 2026
About the Navigator Framework™

Created by Orquestra AI. A working draft, shared for input.

The Navigator Framework™ was created by Orquestra AI in 2025–2026 in response to a structural gap: AI systems capable of more autonomy than most organisations could evidence, let alone defend. The framework was designed to close that gap — not by slowing AI down, but by making Governable Autonomy something that can be measured, implemented, and verified, not just claimed.

The current family is a v0.1 working draft, published for early input rather than a finished standard. Every claim in it is labeled Stable or Provisional so readers know which figures are settled and which are still being piloted. We audit against it, we're building the pilots that will move Provisional figures to Stable, and we publish the working drafts so any organisation can follow along.

Version 0.1 · Working draft · 2026 Not yet released for general adoption Five documents published Trademark application in preparation
How to reference
“Navigator Framework™, Orquestra AI, v0.1 working draft, 2026. orquestraai.io/framework” — or cite the relevant document directly. Status: working draft, shared for early input; trademark application in preparation, not yet registered.

Adopt the framework. Or audit against it.

We run Navigator Audits that score the five capabilities, classify systems into risk classes, and produce a Certified Stage with an Autonomy Contract per system. We also train assessors and governance functions to run the Navigator Capability Model internally.