All articlesAI Engineering

Every check passed. The customer found it anyway.

A model stopped saying anything and kept typing for four thousand words. Three automated checks called it clean - including a second model asked to review the first. Why we now use arithmetic where the textbook uses a judge.

Sep 29, 202612 min read

Every check passed. The customer found it anyway.

One of our agents writes proposal documents. It interviews someone about their business, then produces a document — sections of prose, a few tables, and a closing call to action.

One day the closing section started off perfectly sensibly, and then never stopped:

...cleanly safely smoothly perfectly without delay on track end to end always on schedule smoothly without exception reliably effortlessly cleanly forever directly safely predictably on time cleanly always smoothly securely perfectly end to end without delay across all packages safely smoothly...

That continued for about four thousand words, in a field that normally holds fifteen. It rendered as a dark panel filling several screens of a document someone had waited ten minutes for.

The customer told us. Three separate automated checks had looked at that document and reported that it was fine.

Nothing was broken, which is the problem

No error. No crash. No wrong number. The model simply stopped saying anything and carried on typing. Every word was real, every word was positive, and taken together they meant nothing.

Skim it and it reads like enthusiasm. That quality — fluent, upbeat, empty — is exactly what makes it so hard to catch, and it's why this particular failure is worth writing about.

Three checks, all of them blind

The pipeline wasn't unguarded. Output passed through three checks before it reached anyone.

The first check makes sure a document only contains the kinds of blocks it's allowed to contain. It asks what kind each block is, never what it says. A closing section was allowed there, so it passed.

The second check is our strongest. It pulls every number out of a section and looks for that number in the person's own words or in the researched facts. If a figure appears from nowhere, the section is rejected. We built it after an earlier incident where the model invented a statistic that survived two rounds of review.

It's deterministic, it works, and it was completely useless here — for one simple reason. A wall of adverbs contains no numbers. The check looked at the block, found nothing it was built to examine, and reported clean.

It's a smoke alarm that senses heat. A good alarm. It just can't see this particular fire.

The third check is the one this article is really about. It's a second language model, asked to read the section and fix what's wrong.

We had asked a language model to notice that another language model had stopped saying anything. It read the wall of confident, positive adverbs and approved it.

Where we part company with the textbook

There's a well-argued book on designing systems around language models that lays out the standard approach to this problem, and it's worth engaging with seriously because most teams are following some version of it.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (golden datasets, LLM-as-a-Judge, quantitative evaluation testing); Ch. 6 (validation and evaluation suite, human-in-the-loop) Open source (opens in a new tab)

Its recommendation runs roughly: keep a "golden dataset" of 50–100 representative inputs with ideal outputs; use a second, powerful model as a judge to score your system's answers against those ideals on accuracy, groundedness and tone; turn those scores into a percentage; and block the deployment in your pipeline if the percentage drops below a threshold.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (golden datasets, LLM-as-a-Judge, quantitative evaluation testing); Ch. 6 (validation and evaluation suite, human-in-the-loop) Open source (opens in a new tab)

Most of that we agree with, and we do it. Golden datasets are genuinely useful. Blocking a deploy on a measured drop is the right instinct. The book is also careful to pair the judge with human review of low-confidence cases, which matters.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (golden datasets, LLM-as-a-Judge, quantitative evaluation testing); Ch. 6 (validation and evaluation suite, human-in-the-loop) Open source (opens in a new tab)

But there's a gap in the middle of it, and our broken document walked straight through it.

A judge model is a language model. The output that most needs catching — fluent, positive, content-free — is precisely the output a language model rates highly.

Our third check was an LLM judge. It was pointed at exactly the failure it was supposed to catch, and it waved it through, because from the inside a wall of positive adverbs looks like a nicely written paragraph. Asking a second model to spot the first model's degeneracy is asking the same faulty instrument to check its own reading.

So our version of the rule is narrower, and we'd put it this way: use arithmetic for failures that have a measurable shape, and keep the judge model for questions that genuinely need judgement. Is this answer the right tone for an angry customer? That needs a judge. Has this text stopped containing information? That's arithmetic, and arithmetic doesn't get charmed.

Our first attempt at the arithmetic was wrong too

This is the part worth carrying somewhere else, because we got it wrong in an instructive way.

The standard way to detect repetitive text is to look at every run of three words — every "cleanly safely smoothly", every "safely smoothly perfectly" — and count how many of those runs are new. Real writing produces mostly new ones. A stuck record produces the same ones over and over.

We wrote that check, tested it against the customer's actual broken text, and the test failed. The detector said the broken document was fine.

The reason is the shape of the loop. It wasn't a stuck record repeating one phrase. It was about twelve adverbs being shuffled. Shuffle twelve words and nearly every run of three is, technically, new. The broken text scored 0.97 on that test. Good writing scores about the same.

We had built a detector for a stuck record and been handed a drunk thesaurus.

So we stopped guessing and measured

Four samples: the real broken text, an artificial loop of the same character, ordinary prose, and — deliberately — repetitive technical writing, the kind that says "record" and "status" nine times in a paragraph because that's what the thing is called. That last one is the hostile control. Any detector that flags it will start deleting perfectly good sections.

Two things came out of it, and the first one killed an entire family of detectors.

Length is what separates them, not repetition. At 200 words, a loop and an ordinary technical paragraph are indistinguishable: 0.49 against 0.51 distinct words. Any threshold between those is a coin toss. The signal only appears over longer stretches, for a reason that's obvious once you see it — real writing keeps picking up new vocabulary, and a loop stops. A stuck model isn't introducing new nouns.

Measure a window, not the whole document. The broken block opened with two perfectly good sentences. Average across the whole passage and that good opening pulls the score up. Measure a sliding 120-word window instead and the collapse is impossible to miss. A building inspector walks room to room; they don't average the temperature of the building and call it safe.

SampleWordsDistinct wordsThree-word runsWorst 120-word window
The real loop, full length4,0000.0040.340.14
The same loop, first 200 words2060.490.970.44
Repetitive technical writing780.510.88—
Hand-written prose (our own source comments)1,8380.340.950.58

The last row is the one that matters. Nearly two thousand words of genuine English, written by hand, scoring 0.58 at its worst window against the loop's 0.14. Nothing sits between them. A threshold of 0.35 has roughly 2.4× of room in both directions — not a number somebody picked because it felt about right.

The detector we shipped

goquality/degenerate.go
// Three tests, cheapest first. No model is involved anywhere, on purpose: a
// reviewing model is what failed here, and a loop is the exact input a
// reviewing model is worst at judging.
func looksDegenerate(text string) (reason string, ok bool) {
    words := strings.Fields(text)

    // 1. A phrase repeated six times back to back. Catches a stuck record at
    //    any length, including short ones.
    if phrase, n := longestImmediateRepeat(words); n >= 6 {
        return fmt.Sprintf("phrase %q repeated %d times", phrase, n), true
    }

    // Below 120 words the remaining tests cannot tell a loop from an ordinary
    // technical paragraph, so we do not run them. A detector that fires on
    // real writing deletes good work, which is worse than the bug.
    if len(words) < 120 {
        return "", false
    }

    // 2. The main test: the worst 120-word stretch, by share of distinct words.
    if worst := worstWindowDistinctRatio(words, 120); worst < 0.35 {
        return fmt.Sprintf("distinct-word ratio %.2f in worst window", worst), true
    }

    // 3. A backstop for a loop that shuffles widely enough to keep any single
    //    window looking varied.
    if len(words) >= 400 && trigramRatio(words) < 0.50 {
        return "repetitive across the whole passage", true
    }
    return "", false
}

Note the floor at 120 words. Below that, only the first test runs, because the measurements showed the other two can't tell a loop from a technical paragraph at that length. A detector that fires on legitimate writing would start deleting sound sections — a worse failure than the one we're fixing.

It walks every text field of every block automatically rather than naming the fields it knows about, so a new kind of block is covered the day it exists, not the day someone remembers to add it.

When it fires on a section, the section is written again once, more conservatively. If it loops a second time, the block is dropped and the rest of the section is kept — a document missing its closing line reads fine, while one containing the wall tells the customer nobody read it. And when the text being rewritten already exists, the whole operation is refused rather than saved: a bad rewrite must never overwrite good writing.

Seven tests ship with it. Two of them exist purely to catch false positives — the hostile technical paragraph, and our own source comments as a large body of hand-written English. If someone later tightens the thresholds, those fail loudly before the change can start deleting real sections.

One honest caveat

We added output length limits in the same change, and it would be tempting to describe those as part of the fix. They wouldn't have caught this. At roughly twelve thousand tokens the loop sat comfortably under any limit we'd plausibly have set.

The limits cap the bill. The detector is what keeps the document clean. Being clear about which control is doing the work matters, because the next engineer will otherwise trust a limit that was never load-bearing.

What to take from this

Three things, in order of how widely they apply.

Ask what each of your checks cannot see. Our number-checking audit is genuinely good, and it was useless here for one structural reason: it examines numbers, and this failure contained none. Every check has a shape, and failures find the gaps between them. This isn't an AI problem — it's the same reason a test suite passes while production breaks, because the assertion checked the status code and not the body.

Don't verify something with the same mechanism that produced it. A model checking a model is the same instinct as letting a service decide what goes in its own audit log. For failures with a measurable shape, use arithmetic. For questions that need taste, use the judge — and know which kind you have.

Measure before choosing a threshold. Our first detector was reasonable, plausible and wrong. Twenty minutes of measurement produced a 2.4× separation and a number we can defend to anyone who asks. Constants chosen by taste are the ones that get quietly tuned until they stop firing.

The loop reached a customer because every automated check reported success. What makes us confident about the fix isn't that a new check now reports success — it's that we measured the gap between a loop and real writing and found it wide. That distinction is the whole point, and it's the same one behind every other control we've written about here, from owning state by key to designs that make injection survivable: put the guarantee somewhere it holds, rather than somewhere it merely usually works.

Green checkmarks are not evidence. Margins are.

Sources

  1. Sampriti Mitra, Packt PublishingJun 1, 2026

    System Design for the LLM Era: Patterns and principles for production-grade AI architecture (opens in a new tab)

    Ch. 2 (golden datasets, LLM-as-a-Judge, quantitative evaluation testing); Ch. 6 (validation and evaluation suite, human-in-the-loop)

Thinking about an agent like this for your team?

Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.