Never Let an Agent Grade Its Own Homework

Oct 6, 20267 mins read
JJ Segado | VP, AI Solutions, Argano
At Argano, JJ leads the strategy and delivery of AI solutions, combining deep technical expertise with delivery leadership to help clients turn complex business challenges into measurable outcomes. As VP of Delivery, he brings more than 20 years of experience leading technology teams and enterprise programs.

Part 3 of 6 in a series about Harness Engineering.  Previous articles:
Part 2: The Model Should Never Touch the Tool
Part 1: Harness Engineering

Verification loops, and why the maker can't be the checker  

There is a moment in most agent demos that everyone has learned to skip past. The agent finishes a task, and then — because someone thoughtfully added the step — it reviews its own work and pronounces it good. 

It almost always pronounces it good. 

This is not a quirk of any particular product, and it is not a matter of models being insufficiently self-critical. It is a measurable, well-documented property of language models used as evaluators, and it has a name: self-preference bias. Models systematically rate their own outputs differently from how independent judges rate them, and across studies the effect is large, consistent, and present in essentially every model tested. 

Why the bias exists 

The mechanism is more interesting than a simple story about vanity. Research into LLM-as-judge behavior has found that models assign lower perplexity to their own outputs — that is, their own text is more predictable to them — and that this familiarity correlates with higher scores. The model is not congratulating itself. It is mistaking recognition for quality. 

Even more telling: studies have found that attribution labels alone are sufficient to shift scores. Tell a model that a piece of text is its own and the score inflates. Tell it the same text came from somewhere else and the score deflates. The judgment moves without the artifact moving at all. 

An evaluation that changes when you change the label is not an evaluation. It is a preference. 

Set aside the research for a moment and just consider the structure. You have asked one process to produce an answer under a set of assumptions, and then asked the same process, holding the same assumptions, to determine whether the answer is correct. If the assumption was wrong, both steps are wrong in the same direction, and the review step adds nothing except a false signal of diligence. That is the real cost — not that self-review fails to catch errors, but that it produces a reassuring artifact saying it looked. 

Separating the maker from the checker 

Every production harness worth the name enforces a structural separation between the thing that produces work and the thing that evaluates it. The separation has to be structural — a separate invocation with its own context, its own instructions, and no visibility into the maker's reasoning — because a checker that inherits the maker's context inherits the maker's blind spots. 

In practice this means the verifier should receive the specification and the artifact, and not the transcript. It should not know what the maker was trying to do, only what it was supposed to do. The gap between those two things is where most defects live. 

The loop itself is simple, and its simplicity is a feature: 

  1. The agent produces the work — against a stated specification. 
  2. An independent verifier checks it — against that same specification, with no access to the maker's reasoning. 
  3. On pass, it ships — and the run is recorded. 
  4. On fail, it rejects with a reason — naming the specific flaw, and the loop runs again against that flaw. 

The reject-with-reason step is doing more work than it appears to. A verifier that returns 'this is not good enough' produces a rewrite, which is expensive and often changes things that were fine. A verifier that returns 'the error handling in the payment path does not cover a timeout' produces a fix. The specificity of the rejection determines the cost of the correction. 

Prefer checks that cannot be argued with 

Here is the part that gets lost in the excitement about LLM-as-judge: the best verifier is usually not a model at all. 

If the work product is code, the strongest verification available is the test suite, the type checker, the linter, and the build. These are deterministic. They do not have self-preference bias, they do not have an off day, they cost effectively nothing to run, and they cannot be persuaded. Run them first, and reserve model-based verification for the questions they cannot answer: does this actually address what was asked, does it stay inside the intended scope, does it introduce something the specification forbids. 

The same principle generalizes beyond code. If an agent produces a financial summary, the totals should reconcile against the source before any model reads a word of it. If it produces a document with citations, every citation should resolve. If it fills a structured record, the schema should validate. Deterministic checks are cheap, and every class of error they catch is a class of error you never have to reason about probabilistically. 

There is one anti-pattern worth naming explicitly, because it is common and it is corrosive: an agent that responds to a failing test by modifying the test. It is a locally rational move — the instruction was to make the tests pass — and it destroys the only reliable signal in the system. The verifier has to treat changes to the verification mechanism as a category of change requiring separate approval, or the whole structure quietly hollows out. 

The one place worth paying for the strong model 

Verification is where the economics of a harness invert. Across most of an agent's work — triage, classification, formatting, mechanical edits — a fast, cheap model performs at parity with an expensive one, and routing those subtasks to the cheap tier is most of how teams control cost. We will cover that in detail in the next article but one. 

Verification is the exception. A verifier that misses defects is worse than no verifier, because it converts an unknown risk into a false assurance, and the downstream cost of a defect that reaches a user dwarfs the token cost of catching it. If there is one place in the harness to spend the strongest model available, this is it. 

What verification will not do for you 

Three limits, stated plainly. 

A verifier is only as good as the specification it checks against. If the task was vaguely defined, the verifier has nothing to hold the work up against and will fall back on generic quality judgments, which is close to where we started. Most disappointing verification loops are actually disappointing specifications. 

Verifiers can be wrong in both directions. False rejections are the expensive kind — they burn tokens and wall-clock time on fixes to things that were never broken — which is why the loop needs a retry ceiling and an escalation path to a human, rather than permission to keep going until it convinces itself. Any loop that can run indefinitely eventually will. 

And verification adds latency and cost to every single run, which means it is not free and should not be applied uniformly. Match the rigor to the blast radius. A draft that a human will read before anything happens needs less checking than a change that deploys itself. 

The test to run this week 

If you have an agent in production and you want to know whether this article applies to you, there is a cheap experiment. Take a sample of work the agent approved of its own accord — fifty items is plenty — and have someone qualified review it independently. Not a second agent with the same instructions. A person, or a genuinely separate model with a separate specification. 

The gap between those two approval rates is your defect rate. In our experience it is reliably larger than teams expect, and the number itself tends to end the debate about whether to build this layer faster than any argument does. 

Next in this series: why bigger context windows have not solved the memory problem, and what the agent forgets between turn one and turn forty. 

This is the third article in a series on harness engineering. Part 2 covered tool orchestration and guardrails. Part 4 covers context and memory. ​