Chapter 29. Measuring Whether It Works
What to instrument, what the numbers mean, and which measures are research constructs, not operational metrics.
29.1 Instrument from day one
Nearly every threshold in IRGF is a reasoned default with no empirical calibration. The only route from asserted numbers to calibrated ones is your own data, and it can only be collected prospectively.
Six measures are worth capturing from the first month.
Table 76. Instrument from day one
| Measure | Definition | Why |
|---|---|---|
| Classification effort | Time from S2 start to countersigned score | The dominant recurring cost |
| Gate lead time | Elapsed time per gate, by tier | Detects bottlenecks before routing-around begins |
| Evidence production time | Effort to assemble gate evidence, by tier | Tests whether proportionality is real |
| Tier distribution | Share of estate at each tier | Compared against incident distribution, detects systemic mis-tiering |
| Countersignature adjustment rate | Share of scores changed on countersignature | Tests whether the control is doing work |
| Automated evidence share | Proportion of gate evidence captured automatically | The most honest indicator of falling burden |
29.2 Reading the numbers
Countersignature adjustment rate is the most informative single number. Near zero means either scoring is excellent or countersignature is a rubber stamp; near half means anchors are unclear. Something in the range of one in six to one in four adjusted suggests a control that is working without being adversarial.
Tier distribution against incident distribution is the systemic check. If most of the estate is Tier 1 and most incidents occur in Tier 1 systems, the classification model does not describe your organization.
Gate lead time by tier should diverge sharply. Convergent lead times mean tiering is not changing behavior.
Aged drift backlog is the clearest measure of whether governance keeps pace with delivery. A backlog growing faster than the estate means capacity is insufficient or alert quality is poor.
29.3 The value scorecard
Five metric families, deliberately not composited (Chapter 23, §23.5). Report them together and refuse to average them.
29.4 What you cannot measure yet
Two constructs proposed in the research are not operational metrics: Risk Mitigation Effectiveness and Perceived Governance Burden. Both are research constructs awaiting validated measurement instruments. Neither should appear on a dashboard as though it were measurable today.
The honest position is that IRGF’s effectiveness has not been demonstrated anywhere, and your own instrumentation is the first evidence that will exist. [Practice recommendation] Retain your baseline data. The most valuable contribution an early adopter can make to this framework is a calibrated account of what it actually cost and what it actually caught.