Judgment Zones — Method and Evidence

This page is the back of the model. Judgment Zones says what the model is for and how to start; this one says what is measured, who answers, where each instrument comes from, and what observation would show the model to be wrong. It is written for the consultant, the works council and the sceptic.

What is measured

  • Adaptive potentials, AI edition — perceive, respond, monitor, learn, anticipate. Answered on a bipolar scale from −3 (AI erodes this here) to +3 (AI amplifies it), with 0 meaning no effect yet, so a unit at the very beginning can still answer. Wording adapted from Hollnagel’s Resilience Assessment Grid, out of its safety context and into an operational and market one.
  • Judgment and agency — decision latitude (items after Karasek, rewritten for AI-mediated work), evaluation skill, judgment-broker roles, and the experience pipeline, including whether people deliberately work without AI to keep the skill. These four are the vertical axis.
  • Information culture — the six Westrum items as operationalised by DORA, each with an AI-specific twin. Reported as a gate condition with its own hypothesis, not folded into the profile, so that it can be tested and, if it fails, removed without touching the rest.
  • Entrenchment location — exploration, codification, propagation, obliteration, read from present/absent signal checks per unit and per core workflow. A unit can show signals from two stages; the report shows both. Recorded as a position, never scored.
  • Delegation coherence — one card per decision class: stakes, the stage of automation (acquire, analyse, decide, act, after Parasuraman) and the depth at each, exception rate, who holds override authority, when an override last happened, and what judgment is available to exercise it (Dreyfus level). The card flags incoherence when depth on “decide” or “act” exceeds the judgment available, or when the exception rate is unknown.
  • Viability functions — operations, coordination, resource bargaining, audit, intelligence, identity and policy, plus recursion, channel capacity and the algedonic channel, from Beer’s Viable System Model. For each, one reading: has AI strengthened it, bypassed it, or hollowed it out. Pérez Ríos’s pathology names are used as prompts.
  • Governance — govern, map, measure and manage, taken from the NIST AI Risk Management Framework and recorded as present or absent practices rather than levels: decision classes defined, exception route named, incident log, documented override authority, rehearsed kill-switch, model and supplier register, data-use rule, review cadence. In the depth tier these map to ISO 42001 clauses.

Deliberately not measured: data readiness, technology stack, number of use cases, budget share. They are recorded as context so the research can control for them, and they never appear on a unit’s profile. Counting them is how a model ends up rewarding rollout.

Who answers, and at what level

The business unit is the primary unit of analysis; the core workflow or decision class is the secondary one, because delegation depth and exception handling are only observable there. Team and individual are sampling levels. The enterprise gets a map of its units and an identity question — never an average, never a level.

Whether a team is a sampling level or a unit is not settled by the org chart but by a recursion test: a team that faces an environment of its own — its own customers or counterpart, its own decisions about how to respond — is a viable unit at its own level and gets its own profile. In a small firm with one business unit, the teams or workflows that pass that test are the units.

Decision latitude and evaluation skill are answered by individuals, anonymously. They are never reported individually and only reported at all once at least five people in the unit have answered. This is a design rule, not a courtesy, and it is the answer to the question a works council will ask.

One item set, four layers

The layers differ in who answers, how long it takes and what is reported — not in what the model claims. That is what keeps the tiers from contradicting each other.

  1. Location. Present/absent checks place the unit in a stage or a range of two; the judgment-and-agency bars place it on the vertical axis. Together they give the zone. The zone picture is a communication device with two falsifiable predictions attached — it is not the diagnostic.
  2. Core profile. In v0.1: 28 items answered by the unit’s management team, 18 anonymous member items, and three workflow cards per unit. Output is a profile with polarity, four judgment-and-agency bars, culture gate status, entrenchment location and up to three probes with stop conditions. Fewer than three probes is a valid result. No score.
  3. Viability diagnostic. Interviews along the viability functions, document review of AI decisions and incidents, delegation cards for the relevant decision classes, and a governance record that can be handed to an ISO 42001 implementation.
  4. Research spine. A fixed core — identical for every client — is pooled and fitted against the outcome metrics with every data round. Capabilities that show no link over two rounds leave the core. Anything company-specific stays with the client. Every respondent is told that anonymised core data feeds the model, and a client can opt out.

What the model is tested against

Seven outcome metrics, per unit, at baseline and every six months: decision latency for routine decisions and, separately, for non-routine ones; the probe-to-scale ratio, with the share of probes stopped for a documented reason; governance violations and near-misses, counted separately because near-misses are a positive signal that someone is watching; output per judgment FTE; the experience pipeline; and the override rate together with override quality — how many overrides were later judged correct.

What the model is not tested against: employee satisfaction on its own, enterprise revenue growth, or any vendor-published productivity figure. The first is too far from the mechanism, the second too confounded, the third fails our own evidence rule.

What would make the model wrong

Every structural claim is published with the observation that would refute it and the point at which we expect to have data. Eleven such targets exist in the model draft; these five are the ones that would cost the most.

ClaimRefuted byWhat we would give upExpected
Hollow units degrade: high entrenchment with low decision capacity produces rising non-routine decision latency, rising governance violations and a weakening experience pipeline, relative to Augmented unitsHollow units matching Augmented units on all three across two rounds, with at least 20 units and five per zoneDecision capacity as a predictor; the zone picture becomes description only~18 months after first client data
Pre-AI decision capacity predicts the zone a unit lands in after adoptionStarting decision capacity unrelated to the zone reached, across 15 units with a pre-AI baseline and two rounds afterThe central claim — that decision capacity is a lever and not merely a description2–3 years
The four judgment-and-agency dimensions cohere into one axisNo common pattern across the first 15 unitsThe computed vertical axis; zone placement becomes a facilitated judgment instead of a ruleFirst 15 units
The two scales agree: at re-measurement, the absolute reading and the unit’s own “AI eroded / amplified this” attribution point the same waySystematic divergence across 15 units measured on bothThe bipolar scale — it would be measuring belief about AI rather than change, which would make every erosion finding a self-attribution~18 months
The self-service tier is answerable: a management team completes it in under 15 minutes with fewer than 10 % “cannot answer”Median time above 15 minutes, or “cannot answer” above 10 % on any dimensionThe self-service claim; that dimension moves to the consultant-led tiersFirst 10 pilot units

Evidence discipline

Vendor and analyst figures are labelled as such and never carry an argument. Claims from our own data name the sample size. Borrowed instruments cite their source and the context in which they were validated. Where a validated instrument exists, its logic is reused and credited, the wording is rewritten for the AI context, and the original items are kept alongside wherever comparability matters — which also means the rewritten twins are themselves unvalidated until the pilot fields both.

Where it comes from

Beer’s Viable System Model and Pérez Ríos’s pathologies; Hollnagel’s Resilience Assessment Grid; Westrum’s typology as operationalised by DORA; Karasek on decision latitude; Parasuraman on the stages of automation; the Dreyfus levels of skill acquisition; the NIST AI Risk Management Framework and ISO 42001 for governance; and the entrenchment stages, the outcome metrics and the reading of judgment as the scarce resource from The Tautai Principle. Two systematic reviews of the AI maturity model literature supply the negative case: weak theoretical grounding, largely descriptive use, and no longitudinal evidence that moving up a level improves anything.

Status

Version 0.1, entering pilot use. Nothing on this page has been validated against client data. The item counts, whether there are five potentials or four, and the scale itself are open and will change with the first ten pilot units. Where a figure appears here, it is the v0.1 figure and it is dated by this page.


Back to Judgment Zones.

Scroll to Top