← Back to Articles

An AI Reviewer Without A Rubric Just Flatters You

Strong language models agree with human preferences about which of two answers is better more than 80% of the time (Zheng et al., 2023) — roughly the rate at which two humans agree. That is a real number, and it may be the most dangerous number in AI evaluation, because of what the same work names beside it: position bias, verbosity bias, and self-enhancement bias. The judge leans toward the answer that came first, the answer that ran longer, and the one that sounds like something it would have written.

Ask an open question and a different failure arrives: assistants tuned on human preference data drift toward the answer the user already believes (Sharma et al., 2023). An underspecified question earns a diplomatic answer, and you cannot ship on diplomacy.

GreatFeedback.ai is the system ML LABS built against that gap, and it is live: it scores product feedback against named criteria and rates the reviewer as well. Sixty built-in personas critique work across ten orthogonal lenses before launch; a widget on the shipped page collects reactions from real visitors after. ML LABS builds it, and ML LABS runs it.

Omar Trejo points to a criterion row while an engineering lead writes a score
Named criteria, and a score for the reviewer.

Score Dimensions, Not Vibes

The first design decision is that no reviewer is ever asked for an opinion. Work is scored against named dimensions, each carrying an explicit rubric, and a reviewer forced to rate accuracy, clarity, and persuasiveness as separate axes cannot return a blended, noncommittal shrug. The score is worth exactly what the rubric behind it is worth — which is why the rubric, not the model and not the prompt, is where the engineering effort goes.

  • Orthogonal axes stop one strong quality from masking a weak one in a blended number.
  • A defined rubric for each dimension hands the reviewer a fixed standard, and that fixed standard is what converts an opinion into a repeatable measurement.
  • Custom dimensions let the work carry its own criteria: the ten standard lenses ship with the system, but the standard for this particular piece of work does not.

Sixty personas across ten lenses is not a feature count. It is a refusal to average. A skeptical buyer, a domain expert, and a first-time visitor reading the same page disagree in structured ways, and a mean taken over their scores destroys the only information worth having. Writing the standard down before measuring against it (Google, 2016) is the discipline that makes a service objective more than a dashboard, and it is the same discipline here: the target exists before the number, or the number means whatever the model wants it to mean.

A single reviewer returns an average. A panel built to disagree returns a map of where the work is contested.

Critique From Both Sides Of Launch

Pre-launch and post-launch feedback answer different questions, and GreatFeedback ships both since neither substitutes for the other. The persona panel says what is likely wrong before anyone has seen the work, instantly, and cheaply enough to run on a draft. The widget says what is actually wrong, from people under no obligation to be kind. Synthetic critique has speed and no ground truth; real reaction has ground truth and no speed.

The artifact worth having is the gap between them. A predicted objection nobody raised, and a real objection no persona anticipated, are both calibration data on the critic — and a review system that cannot measure its own miss rate should not gate anything. This is the same split as the difference between output metrics and outcome metrics.

That is why human ratings of the AI's feedback feed back into persona quality. The critic is itself under review, and the correction signal is a person marking a note as useful or as noise. An unrated critic is an unmonitored model: the finding that ML systems accrue hidden debt wherever monitoring and feedback are missing (NeurIPS, 2015) lands on the reviewer exactly as it lands on the thing being reviewed. Its rubric interpretation drifts, its calibration on your material decays, and the output reads as confident the entire way down.

Where Rubrics Run Out

Rubrics have a boundary, and the score does not stop at it. Genuine originality, whether a joke lands, whether a design has any soul — these resist decomposition into scored axes, and forcing them into axes produces a confident measurement of the wrong thing.

The honest architecture uses the panel for everything a rubric covers well — clarity, correctness, structure, coverage, whether a claim survives a skeptic — and routes the irreducibly aesthetic judgment to a human who owns it and signs for it. A rubric is a floor under quality, never a ceiling on it. Sold as the ceiling, a scoring system eventually gates out the work that was worth shipping precisely because it did not fit an axis.

First Steps

  1. Write the dimensions and their rubrics before writing a prompt. If you cannot say what a good score means per axis, the reviewer will decide for you, generously.
  2. Replace the single reviewer with a panel across lenses that genuinely differ, then read the disagreements before the consensus — the split is the finding.
  3. Rate the critic. Capture a human judgment on each piece of feedback, and route a real signal from after launch back against the prediction made before launch.

Build The Rubric, Then The Reviewer

The critique call is the cheap half of a review system. Ten lenses that do not collapse together, sixty personas coherent under them, custom dimensions the work defines, and a human rating loop correcting it all — that is the engineering half, the half a prompt cannot buy. The model underneath stays swappable, bound to a standard it did not choose.

This becomes a system rather than a prompt when review gates a real decision — a ship, a spend, a release. Guidance on placing explicit checkpoints in front of consequential actions (Anthropic, 2024) applies to review as squarely as to execution, the same instinct that keeps the model on the surfaces while a deterministic core owns anything irreversible.

Anyone shipping AI-generated work at volume with no rubric behind the review is standing exactly where this pays. The next step is one contained workflow taken all the way to production: an engineering partnership writes the acceptance targets into the SOW before work starts and is kept running by the person who built it. The rubric and the correction loop are the deliverable. The model is the part you are free to change your mind about later.

References

  1. Zheng, L., et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv, 2023.
  2. Sharma, M., et al. Towards Understanding Sycophancy in Language Models. arXiv, 2023.
  3. Sculley, D., et al. Hidden Technical Debt in Machine Learning Systems. NeurIPS, 2015.
  4. Anthropic. Building Effective Agents. Anthropic Engineering, 2024.
  5. Google. Service Level Objectives. Site Reliability Engineering, 2016.

NEXT · TO PRODUCTION

Find the real blocker.

What is slowing delivery, and the fastest path through.

15 minutes · no charge · with Omar