← Back to Articles

Your Delivery Metrics Look Healthy While Risk Grows

Research on temporal degradation in deployed models (Scientific Reports, 2022) tested 128 model-dataset pairs across healthcare, weather, finance, and transportation, and found 91% of them degrading over time. None of those degradations would have moved a DORA metric. Deployment frequency, lead time, change failure rate, and mean time to recovery would have stayed where they were while the systems got worse in production.

DORA earned its place as the default engineering scoreboard, and the four metrics still measure something real. But they rest on an unwritten assumption: that the humans shipping the code understand it, review it, operate it, and can explain it under pressure. When much of a change set is drafted by a coding assistant, glued by an agent, or copy-edited by a model, the deployment record stops being a faithful proxy for what the team knows.

The control that answers this gap is not a new metric; it is a separation of duties, and it has to be structural to hold. In the engineering loop ML LABS runs, the builder and the verifier are drawn from different model families, so the model that wrote a change is never the model that approves it; the verifier reads the spec and the diff cold, with none of the builder's accumulated justification; and the loop has four named stop states, one of which is escalation, the task as written is wrong. None of that produces a number on a delivery metrics dashboard. All of it decides whether the change should have shipped at all.

An engineering director holds a printed code change and asks a senior engineer to explain it at his desk
Shipped on time; nobody can explain the change.

Three Ways DORA Misleads Now

The four metrics measure the surface of delivery: how fast work reaches production, and how the system behaves when a change misbehaves. When changes flow through that nobody can reconstruct, the metrics register the activity but not the comprehension gap.

Legibility Drops While Velocity Holds

AI-assisted code can produce volume indistinguishable from expert output on a dashboard. Pull requests merge, deploys succeed, lead time compresses. What is missing is the mental model that used to be a byproduct of writing the code: the engineer who accepted a suggestion did not necessarily build the understanding that suggestion implies.

Review degrades from the other side. As the volume of plausible-looking diffs rises, review compresses toward a plausibility check: does the output look right, do the tests pass. The substantive judgment, whether this change belongs here and whether it composes with the surrounding system, gets crowded out first, because it costs the most attention.

Hidden Changes Bypass The Pipeline

A growing share of critical change lives outside the surface DORA observes. Vendor toggles, identity provider rules, feature flags, model prompts, agent tool definitions, and infrastructure drift reshape behavior without producing a pull request. None appears in deployment frequency. None counts against change failure rate when it breaks. A dashboard scoped to the repository measures the part of the system that was already under control.

Fast Recovery Masks Brittle Systems

A low mean time to recovery reads as operational strength. It can also be the opposite: the team rolls back aggressively, the critical path returns, the metric records a clean recovery, and the question of what broke is never answered. The rollback ends the incident without ending the condition behind it, and the next expression of that condition arrives against a different hidden surface, where the dashboard again looks healthy.

The recognizable version is the unacknowledged dependency the on-call team was certain could not be the cause. A rollback proves it was, service is restored, and the metric reports a win while staying silent about the comprehension gap that produced the incident. Fast recovery without root-cause clarity is a leading indicator dressed as a lagging one.

Flowchart: AI-assisted delivery raises deployment frequency, so the dashboard looks healthy, while team comprehension drops, so the critical path becomes unexplainable. Both lead to false confidence and then to an outage from a hidden surface.

When the dashboard and the team's actual grip on the system point in opposite directions, the dashboard is the one that is wrong.

Pair DORA With Comprehension

The fix is not to throw DORA out; the pipeline metrics still tell you whether the delivery machinery works. They need a companion frame that observes what they cannot.

Three augmentations carry most of the weight. Every change set carries a short scope note naming its risk profile, the surfaces it touches outside the main repository, and the vendor or identity configuration it depends on. Quarterly metric reviews are paired with two or three real incident narratives from the quarter, root cause written out rather than compressed into a recovery-time number. And the team tests its own legibility on a regular cadence: can members independently explain the critical path in plain language, and where do the explanations diverge. The divergence between them is the measurement.

The moment DORA becomes a scoreboard for grading individuals, the team optimizes for the dashboard instead of the system, and the gap widens in the one direction nobody is watching, which is what a live AI system costs when nobody owns it.

When DORA Alone Is Still Enough

On a team with a single owner of the system, where AI assistance is light and most changes pass through the person who wrote the surrounding code, DORA alone remains a reasonable read of engineering health. The system fits in one head and that same head is on call. Augmentation earns its cost where the change surface is widening faster than the understanding of it, which is the condition AI-assisted delivery creates by default.

Measure Comprehension, Not Velocity

The bar is short enough to hold your own organization against this quarter. The generator of a change is never its verifier, whether both are people or one is a model. Every change carries the surfaces it touches outside the repository. Every incident produces a written root cause, not just a recovery time. And somebody can explain the critical path without opening the code, because a system nobody can explain is a system nobody is operating.

Where the gap is already open and nobody inside the team has the vantage point to see it, that is an ownership problem before it is a metrics problem. Advisory puts an accountable owner on a system we built, with monitoring, drift and cost review, and a written operating brief every month, month to month on 30 days' notice, which is what turns comprehension from an aspiration into somebody's job. The separation of duties that keeps the shipping side honest is covered in why AI coding agents need separate verifiers.

References

  1. Vela, D., Sharp, A., Zhang, R., Nguyen, T., Hoang, A., and Pianykh, O. S. Temporal Quality Degradation in AI Models. Scientific Reports, 2022.

NEXT · TO PRODUCTION

Check your position.

Two minutes. Your main blocker and first move.

15 minutes · no charge · with Omar