This is Part 10 of the ten-part series AI Governance in an AI-Native Software Development Company.
Governance maturity is easy to overstate because documents are easier to see than runtime behavior. A policy library, a committee, a registry or a review process can all be useful, but none of them proves that an agent will be stopped, constrained, attributed or recovered when the system is under pressure.
The practical test is simpler: what can the runtime demonstrate without the team explaining what it intended to do?
This maturity model uses that test. It has six stages, from Experimental to Autopiloted, and each stage is defined by operational evidence rather than aspiration.
Stage 0: Experimental
Agents are already being used, but governance is not yet treated as an organizational problem.
Tools spread through individual adoption. There is no reliable inventory, no clear owner, and no common answer to basic questions such as which agents can reach production data or who is accountable for an agent-authored change.
The move to Stage 1 is mostly organizational. Someone has to name the problem, own it and make AI usage visible enough to discuss.
Stage 1: Ad Hoc
Governance exists as a conversation, not as a system.
A few people know which tools are in use. Rules appear in chat threads, review comments and internal pages. Decisions depend on who happens to be involved. The organization may react well to individual incidents, but it cannot reliably reproduce the same decision across teams.
The next step is to turn tacit knowledge into an inventory and an explicit operating policy.

Stage 2: Documented
The organization can describe what it has and what it expects.
Agents, skills and tools are inventoried. Policies are written and owned. Teams can point to standards for data access, approvals, change management and acceptable use.
This is progress, but it creates the most common false ceiling in AI governance: documentation can become more mature while runtime behavior remains unchanged.
The transition to Stage 3 is the expensive one. Policies have to become machine-evaluated artifacts, and enforcement has to exist at the points where behavior can actually change.
Stage 3: Enforced
Policy is now part of execution.
Contracts and policy artifacts are versioned. Authoring, CI, deployment and runtime contain enforceable gates. High-value actions can be refused or escalated. Containment is not a paragraph in a runbook; kill, scope and recovery mechanisms are exercised. Important changes can be traced from intent to production effect.
At this stage the question changes from “do we have controls?” to “are the controls working well?”
Stage 4: Measured
The governance system is observable on its own terms.
Teams track more than adoption. They measure blocks, false blocks, false allows, exception lifetimes, containment time, incident attribution and the cost of governed actions. The same audit question should produce the same evidence next quarter, not a new manual reconstruction.
Stage 4 is a valid destination. Not every organization needs to automate routine governance decisions further. If the fleet is stable and the blast radius is bounded, human interpretation of reliable measurements may be the right design.
Stage 5: Autopiloted
Routine governance decisions are handled automatically against signed, versioned artifacts. Humans focus on exceptions, novel risks and policy evolution.
New agents inherit controls by deployment rather than by memory. Exceptions expire automatically. Routine drift can trigger predefined responses. The governance team can cover a larger fleet without growing linearly with it.
There is no useful Stage 6 called “no humans.” Human accountability remains part of the system. Autopilot changes where humans intervene; it does not remove responsibility.
The documentation ceiling and stage inflation
The easiest maturity mistake is to evaluate the program by visible activity.
A team can have excellent documentation and still be Stage 2 if policy violations are not actually blocked. It can have a dashboard and still be below Stage 4 if the dashboard measures only how many agents were deployed.
A useful self-assessment rule is strict: your stage is the highest stage for which the critical capabilities can be demonstrated end to end.
One missing control on a high-risk path matters more than five polished controls on low-risk paths.
Maturity should be proportional to risk
The top of the ladder is not the goal for everyone.
The correct target depends on blast radius. An internal assistant that cannot modify production or access sensitive data does not need the same control depth as an agent that can move money, change customer-facing systems or act across regulated data.
That leads to a better planning question:
What is the lowest maturity stage that keeps this specific risk defensible?
Over-governing low-risk surfaces wastes engineering effort. Under-governing high-risk surfaces creates exposure. A mature program allocates controls according to consequence, not according to the prestige of a stage number.
Measure adoption, enforcement and outcomes together
A governance program needs three layers of metrics.
Adoption metrics show reach: registered agents, teams onboarded, skills published and policy coverage.
Enforcement metrics show whether controls execute: policy evaluations, block rates, false-block and false-allow rates, exception expiry, containment engagements and time to contain.
Outcome metrics show whether the program changes the business risk: incident rate and severity, audit-query latency, cost per governed action, recurring failure themes and the time from an agent-authored change to a production effect with provenance intact.
Any one layer by itself can mislead. Adoption without enforcement is documentation. Enforcement without outcomes is process. The three together describe whether governance is actually improving the operating system.
A DORA-like discipline for agents
Software delivery improved when teams learned to read throughput and stability together. Agent governance needs the same discipline.
Do not celebrate more agent-authored changes without also reading failure and recovery signals. Do not celebrate faster execution without knowing whether provenance, containment and policy compliance stayed intact.
Throughput without stability is not maturity. It is faster exposure.

Runtime Verification Probes
Questionnaires encourage optimistic self-assessment. Probes ask the runtime for evidence.
Try a few concrete tests:
- Provenance probe: choose a recent production change and trace the agent, model, policy version and authorizing intent.
- Containment probe: stop a real agent in a safe environment and prove that the stop took effect within the expected bound.
- Calibration probe: show the recent trend for false blocks and false allows by policy.
- Honesty probe: show a governance metric that was uncomfortable and the action taken because of it.
- Autopilot probe: identify a routine governance decision completed automatically against signed artifacts, with evidence.
A stage claim that cannot survive its corresponding probe is an aspiration, not a runtime capability.

How to use the model on Monday
Do not turn the model into another transformation program before using it.
- Run the probes and place the organization at the stage the runtime can prove.
- Identify the highest-risk missing capability at that stage.
- Pick one concrete control that closes the gap.
- Assign one owner, one deadline and one demonstrable completion artifact.
- Choose an enforcement or outcome metric that will show whether the change worked.
- Schedule the next-stage probe and run it again at the end of the cycle.
The goal is not to climb the ladder for its own sake. The goal is to make the next governance decision more evidenced, more repeatable and more defensible.
Closing the series
Across ten articles, the architecture is consistent.
Identity tells us what is acting. Contracts define what it may do. Executable policy turns intent into runtime decisions. Containment limits blast radius. Traceability connects intent to effect. Change management treats capability changes as production changes. Budgets bound cost, tokens, capability and time. Measurement shows whether the complete system is working.
Maturity is the final question: how much of that architecture can your runtime prove today?
A governance control that cannot be demonstrated will eventually decay into documentation. A mature AI-native organization keeps forcing the runtime to show that its controls are still real.
Originally published by Reza Arani on Medium in June 2026. Adapted for Aipolix as Part 10 of the AI Governance in an AI-Native Software Development Company series.


Comments
No comments yet. Be the first to comment.
Leave a comment