← Back to Articles

The Control Your Auditor Cannot Re-Perform

There is a step in your close, your revenue process, or your approvals where the control is a person. They check the thing, and they are good at it. The auditor asks for evidence that the check took place for a sample of thirty transactions, and the answer is a mail thread, a spreadsheet with no history, and a recollection. The control may have been performed every time. It cannot be re-performed, which for audit purposes is close to it not existing.

A control an auditor can re-perform has four properties, and none of them is about the software being clever: the evidence is captured at the moment of the action, the reasoning is captured with it, the population is complete, and nobody could have rewritten the record afterwards. Software is good at all four, and bad at them if nobody asked for them at design time. A system built without an evidence path produces a faster process and a weaker control, and the second half surfaces only in fieldwork, when it is expensive.

ML LABS engineered the backend of a medical-device cloud platform that has reached clinical production in two countries and keeps expanding, built for HeartSciences, with AIM Consulting building the frontend. Every design decision in it carried evidentiary weight before it carried engineering weight. What follows is the execution model that made that possible, stated as three control layers. Each one closes a failure that the other two cannot.

  1. Evidence layer that records data assumptions, results, and decision reasoning.
  2. Change layer that governs how models, rules, or workflows are modified over time.
  3. Operating layer for human review, escalation, monitoring, and rollback in production.

Each layer feeds the next, and production feeds new evidence back into the first.

Omar Trejo points to one line of an auditor's approval sample as the controller takes notes
A habit is not a control until it leaves proof.

Evidence Is A System Behavior, Not A Binder

The first question is not "can this work?" It is "what evidence would make this acceptable to operate, and to be examined?" That evidence covers where the data came from, what the system was evaluated against, how it behaves on edge cases, and why somebody accepted the threshold. It is an engineering artifact, not a documentation exercise.

Evidence reconstructed after a build is thinner and harder to defend than evidence designed into the execution path: the facts that make a trail useful, what the state was before a change and why it was made, exist only at the moment of the change. Afterwards you are reconstructing, and that is a story. The AI Risk Management Framework (NIST, 2023) treats governance, traceability, and monitoring as operating requirements rather than as reports.

On the ECG platform, clinical reports move through a lifecycle that is signed, confirmed, and locked, and every modification after signature is tracked with a full audit trail, because the report is a controlled artifact and its history is part of it. Every record reaches a definite terminal state, and that state is queryable. Nothing is "still processing since Tuesday", which turns a completeness check into a database query instead of an act of heroism.

The evidence layer is not a binder you assemble at year end. It is a system behaviour you design before the first line of production code runs.

Three grouped layers in a loop. The evidence layer (data provenance, decision rationale, evaluation results) defines what can change in the change layer (change taxonomy, approval gates, re-validation triggers), which sets how it runs in the operating layer (human review points, escalation paths, rollback triggers), which produces new evidence.

Change Control Before Drift

Models, rules, reference data, and human review thresholds all change, and each is a controlled change. The change model has to exist before the system matters: which changes are pre-approved, which require re-validation, which force a rollback.

On the ECG platform, security work ran inside the delivery cadence: rotating secrets and least-privilege access, sensitive data kept out of logs and errors, input validation on every external interface, and per-organization session isolation with no fallback authentication path. Each finding became a controlled change with its own evidence trail.

Fail-Safe By Default

The instinctive approach to sensitive data is to strip out every field known to be sensitive. That approach fails silently when a new field shows up which the filter does not cover. The stronger pattern reverses the default: define what is permitted to pass, and block everything else. Any new field is held back until it has been classified and approved.

A cardiologist and a clinical systems specialist check an ECG review record at a hospital workstation
Every record must show who changed what.

The Controls That Run Every Day

A system that carries a control needs its behaviour specified as precisely as its decision logic: who approves exceptions, who sees low-confidence output, how incidents escalate, and how it reverts when its behaviour drifts. Without that, "human oversight" is a phrase in a slide, which is the distinction on which who signed off, and what the record shows turns.

Separation between organizations is enforced at every layer as a system invariant rather than an application convention. Each integration point (identity providers, upstream data systems, payment processors) needs its own audit trail, because each is a boundary where responsibility changes hands. SP 800-53 (NIST, 2024) requires organizations to generate, protect, and retain audit records, so every integration boundary is a logging boundary.

First Steps

  1. Name one process and one control inside it. Narrow enough that the evidence path can be designed concretely for every state transition the process contains.
  2. Define the change taxonomy. Which changes are pre-approved, which require re-validation, which trigger a rollback; fail-safe by default for sensitive data.
  3. Write the control loop. Authentication, session management, logging, separation, and rollback triggers, specified before the system becomes operationally important.

Design The Control Path First

The bar is checkable before anyone writes production code, and it is short. Every consequential action produces a durable record naming who acted, what changed, why, and what the state was before and after. Every category of change has a defined path. Every record reaches a terminal state a query can name. And sensitive data flows through an allowlist, so a field nobody has classified cannot leave. A company that cannot point to all four will retrofit them under time pressure, which is the same work at higher cost.

A fifteen-minute call turns one workflow into a plan naming its evidence path, its change taxonomy, and what it would cost to build. The build that follows is one workflow, with the repository (code, intellectual property and full history) transferring to you as it is paid for. Where production data cannot be granted, the work runs against a de-identified or synthetic extract. The full story behind this model is the AI-ECG management platform.

References

  1. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST, 2023.
  2. National Institute of Standards and Technology. Security and Privacy Controls for Information Systems and Organizations (SP 800-53 Rev. 5). NIST, 2024.

NEXT · TO PRODUCTION

Check your position.

Two minutes. Your main blocker and first move.

15 minutes · no charge · with Omar