Judgment Zones
An AI capability model that tells you where each part of your organisation stands — not what grade it gets.
Your units are in different places, and they should be. The one with the most AI is not necessarily the best placed. Judgment Zones locates every unit on two axes: how deeply AI is entrenched in its work, and whether the people acting on AI output can still tell when it is wrong.
One level for the whole company tells you nothing
Almost every AI maturity model in circulation has the same shape, inherited from the Capability Maturity Model: five levels, a set of dimensions scored against them, one company-wide result, and the implicit claim that the top level is everyone’s goal. Two systematic reviews of the field find weak theoretical grounding, largely descriptive use, and no longitudinal evidence that moving up a level improves anything. Just under half of the AI models reviewed assess technology only.
In practice that shape produces three errors.
- It averages away the difference between units. Marketing and production are not in the same place and should not be marched through the same gate. A single company level hides the one thing a leadership team needs to see.
- It counts tool rollout as progress. Access without the discretion and the skill to override is not capability. It is exposure — and a model that scores rollout rewards the failure mode.
- It measures how much AI is used, not whether anyone can tell when it is wrong. The scarce resource is the capacity to filter, validate and direct AI output. No published maturity model measures it as a first-class dimension.
Four zones, two axes
AI entrenchment (horizontal) is how deeply AI is woven into a unit’s routines, talent systems and culture — from first exploration to work that could not be undone. Decision capacity (vertical) is the unit’s ability to filter, validate, override and direct what AI produces: decision latitude, evaluation skill, judgment-broker roles, and where the next people who can do this come from.

Which of your units is in which zone — and do you know, or are you guessing?
For at least one unit the answer is usually a guess. That gap is the reason to measure. Location, not grade: units of the same company sit in different zones at the same time, and the model never averages them into one score.
What you get is a decision basis, not a report
Where to invest next, per unit
Some units need tooling. Some need decision latitude. Some are right to wait. You stop buying one programme for an organisation that is in four different situations.
Delegation matched to judgment
You delegate to AI only as deeply as you can evaluate. Where delegation has run ahead of judgment, the model flags it per decision class — before the exceptions arrive.
Governance from the first pilot
Decision classes, exception routes and an override record exist at exploration stage, sized for a mid-sized firm — and grow into ISO 42001 documentation without a rewrite.
A trend, not a snapshot
The same items are re-taken every six months against outcome metrics — decision latency, probe-to-scale ratio, governance violations and near-misses, override quality — so you learn whether your own moves worked.
Four things a ladder cannot do
- Decision capacity as a first-class axis. Whether the people acting on AI output can tell when it is wrong, are allowed to override it, and where the next people who can do that come from. No published AI maturity model measures this. It is the axis that separates a working rollout from a hollow one.
- One item set behind all tiers. The self-service profile and the two consultant-led tiers read the same items, so the tiers cannot contradict each other — and a unit measured before it adopts anything can be measured again afterwards on the same instrument.
- The model can lose. Every structural claim is published together with the observation that would refute it and the point at which we expect data. A model that cannot fail is a marketing device.
- Evidence discipline. Vendor and analyst figures are labelled and never load-bearing. Borrowed instruments cite their source and their validation context.
The lineage is explicit: viability rather than fitness, from the Viable System Model; entrenchment stages and judgment as the scarce resource, from The Tautai Principle; agency before access, from the research on access without discretion.
Four ways in
Judgment Zones Check
20 minutes · one picture · no preparation
You place your own three to five units on the zone chart, live. Disagreement in the room is the finding. You leave knowing which units you can describe and which ones you have been guessing about.
Judgment Zones Profile
Online · 10–14 min for the management team, 5–7 min per team member · report within a day
A real per-unit profile from your own answers, without a consultant in the room: adaptive potentials, judgment and agency, culture gate, entrenchment location, and up to three probes with stop conditions.
Judgment Zones Potential
For a unit that does not yet use AI · three to four and a half consultant days per unit
A pre-AI baseline that can be measured again later: why there is no AI here and whether the stated reason is the real one, how deep this unit could delegate given the judgment it has today, and what governance must exist before the first experiment.
Judgment Zones Diagnostic
For a unit already using AI in a core workflow · three to five consultant days per unit
Which viability functions AI has strengthened, bypassed or hollowed out; where delegation exceeds judgment, decision class by decision class; and a governance record you can hand to an ISO 42001 implementation.
What none of the four produce: a score, a certificate, a benchmark ranking against your industry, or a roadmap to level five. “Do not start in the next twelve months” is an admissible result.
Who it is for
A good fit
- Leadership teams in mid-sized organisations that are past the first pilots and can no longer say what they actually have
- Internal and external consultants who need a defensible way to talk about AI adoption without selling a five-level ladder
- Managers of a single unit inside a larger group — the model reads at unit and workflow level, no corporate programme required
- Units that have not started at all: they are read as a pre-AI baseline, not as a zero
Not a fit
- Organisations that want a certificate
- Organisations that want a benchmark ranking against their industry
- Organisations that want a roadmap to
“level 5” - Those are available elsewhere, and we will say so.
Where this stands
The model is at version 0.1 and entering pilot use. The zone claims and the links between capability and outcome are stated as hypotheses, each with the observation that would refute it and the point at which we expect data. They have not yet been tested against client data. We say so here rather than after you have bought something.
Start with the guess you cannot defend
Twenty minutes, one picture, your own units on it. If the room disagrees about where a unit sits, you have found the reason to measure. If it does not, you have lost twenty minutes.
How the model is built, what it measures, who answers it, and what observation would refute it: Method and evidence.