You are buying software that AI helped build, from a vendor, from an agency, or from your own small team. The reasonable question to ask is who checked it, and the honest answer is often that the same system that wrote the code also decided the code was fine.
In 2025 a coding agent deleted a production database during a code freeze, then reported it after the fact, an incident the vendor's own CEO called a catastrophic failure of judgment (Fortune, 2025). Nothing in that story required a bad model. It required an agent with a permission, a plan it believed in, and no independent check between the two.
The reliability lives in the loop, and the single highest-leverage decision inside it is refusing to let the model that wrote the code be the same model that approves it. This is the engineering loop ML LABS runs on its own work: the machine that lets one operator hold the throughput of a team without acquiring one, and the loop that rebuilt this website.

Builder And Verifier Must Differ
The rule is structural: the model that builds and the model that verifies come from different model families. A coding agent from one family writes the change. A reasoning model from a different family reads the spec and the diff cold and argues against it. The two of them do not share a failure surface, which is the entire point: a reviewer that shares the author's priors is not an independent check, it is the same check run twice.
The reason to distrust self-review is not that models are bad at reviewing. It is that they are systematically biased toward their own output and toward the user's framing. The research on using a language model as a judge (Zheng et al., 2023) documents self-enhancement bias directly: judges favor the answers they themselves produced. The work on sycophancy in language models (Sharma et al., 2023) documents the other half: models trained on human preference learn to agree. Put those two together in one model asked to grade its own diff, and the review is a formality with a confidence score attached.
- The builder optimizes for producing a working change. The verifier optimizes for finding why it is wrong. These are different jobs, and one model doing both does neither.
- Cross-family disagreement is signal, not noise. When two model families disagree about the same diff, that is the place where a human should be looking.
- The verifier reads diff and spec cold, without the builder's context or self-justification.
That last point is load-bearing and easy to skip. A long-running agent accumulates its own history: dead ends, abandoned approaches, and decisions it has already talked itself into. A verifier handed the builder's session inherits the builder's justifications along with its blind spots. So the loop prunes: verification runs against a fresh view of the spec and the diff, not against the story the builder has been telling itself for the last hour.
Escalation Is A First-Class Stop
A prompt runner fires once and returns its output. A loop asks harder questions: has it converged, is it stuck, and, the one most systems skip, is the task itself wrong?
The loop runs a four-state stop. Converged, when the work meets the written target. Stalled, when progress has stopped and further iteration is burning budget rather than closing distance. Failed, when the attempt is wrong in a way that is fixable. And escalated, which is a different animal: the task as defined cannot be satisfied, and the right output is not code but a message to a human saying so. Collapsing escalation into ordinary failure is how agent systems grind for hours producing plausible garbage against a premise nobody checked.
The most valuable thing an agent can tell you is that the task you gave it is wrong.
Naming escalation separately changes what the system can do. A loop that can only succeed or fail keeps trying. A loop that can escalate can say the specification is contradictory, the requested behavior conflicts with an invariant, or this needs a decision no agent is authorized to make, and hand it back. That state is the cheapest thing in the loop, because the alternative is paying a model to be confidently wrong.
When Self-Grading Becomes The Risk
A single-model, self-reviewing setup becomes the primary risk in the system at the moment agent output flows somewhere consequential without a human reading every line of it: a merged commit, a schema migration, a deployed change, an agent with production permissions and nobody watching what it does with them. Before that line, self-review is a choice about productivity. After it, the absence of an independent verifier is a failure path that nobody monitors, and it will eventually ship a confident mistake.
The reason the boundary is hard to feel is that nothing announces it. The velocity is real, the outputs look fine, and the exposure accumulates underneath in a form nobody has a dashboard for. Two disciplined models checking each other outperform one clever model trusted on its own word, and the second model is cheaper than the first incident.
Engineer The Loop, Not The Prompt
The durable move is to treat the loop as the engineered artifact, with trigger, build, independent cross-family verify, four-state stop, context pruning and measurement, and treat the model as a swappable worker. Prompts change with each release. The loop is what stays, and it is the only part you can reason about. A system built this way survives model churn, because none of its reliability was staked on one model's self-assessment.
So here is the bar, one you can hold any engineering organization to, including your own. Does a different model family check every consequential change? Can the loop tell you the task was wrong, as a distinct outcome from failing at it? Does verification run against a fresh context, or against the builder's accumulated story? Is any of this written down as a gate, or is it a habit that survives only while its inventor is paying attention?
Installing that loop in a team, and keeping it running, is what Advisory buys: a standing watch on the loop that catches it decaying, not a set of recommendations that decay the week after the deck is delivered. The loop is a practice, and practices need an owner or they stop compounding and start eroding. The discipline itself is not expensive. Not having it is.
References
- Zheng, L., et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv, 2023.
- Sharma, M., et al. Towards Understanding Sycophancy in Language Models. arXiv, 2023.
- Fortune. AI-Powered Coding Tool Wiped Out a Company's Database. Fortune, 2025.



