You've had this experience. You ask an AI tool for a logo, a landing page, a business plan. What comes back is fine. The grammar is right, the sections are all there, the colors don't clash. And you feel nothing. You couldn't pick it out of a lineup of ten others, and neither could your investor.

That reaction is worth taking seriously, because it's not a failure of correctness. Nothing is wrong. The work is competent. It's just interchangeable, and interchangeable work doesn't win a customer, close a round, or make anyone remember your company.

We spent months chasing this in our own product, assuming it was a model problem. It wasn't.

The 30-second version

Generation regresses to the mean unless something pulls it away from the mean. The model has seen a million landing pages and produces the center of that distribution, which is by definition unremarkable.

The thing that's supposed to pull it away is evaluation. But most AI systems, ours included until recently, evaluate output in a vacuum: they ask a model "is this good?" without showing it what good looks like. A grader with no reference grades against its own priors, and those priors are the same ones that produced the work. It's marking its own homework.

Fix the reference, not the generator. That single change lifted logos, landing pages, mobile apps, and PDFs at the same time, because they were all failing for the same reason.

Three mechanisms that manufacture "fine"

1. The average of everything is unremarkable

Ask for "a clean, modern brand for a plant care service" and you'll get a sans-serif wordmark, a sage-green palette, and a leaf. Not because the model lacks imagination, but because that is the statistical center of everything labeled "clean, modern, plant." The model answered your question correctly. The question had one obvious answer, and obvious answers are shared.

Distinctive work comes from constraints that push away from the center. "Show me the brand's actual world" is such a constraint. "Clean and modern" isn't a constraint. It's a request for the middle.

2. Judging without exemplars

This is the expensive one, and it hides well.

If you ask a model to score a deck out of ten, it will give you a number, and the number will feel meaningful. It isn't. Without a reference set of genuinely world-class decks in front of it, the model scores against a vague internal notion of "deck-ness," which sits at the same middle as the generator. So it reliably returns sevens and eights for work that a design lead would call unremarkable, and worse, it approves it. The loop terminates. Nothing improves.

Put three genuinely excellent examples beside the candidate and ask "which of these four would you show a client, and why is the candidate not one of them?" and the same model, with the same weights and the same prompt structure, becomes sharply, usefully critical.

3. Templates add; they never subtract

Every generated document accretes. A section for the market, one for the team, one for the roadmap, a risks appendix, a glossary. Each addition is individually defensible and the sum is exhausting.

Actual craft is mostly removal. The consultant's deck that lands is the one where forty slides were cut to twelve. Almost no generative system has a subtraction pass, because subtraction is hard to justify line by line. You can always argue for keeping something. We had to add one deliberately, as its own step, with permission to cut.

What we changed

We stopped letting a model assert facts it can compute. Our brand books used to report accessibility contrast ratios. The numbers were confidently stated, professionally formatted, and invented. Contrast is arithmetic. You take two colors and calculate. Now we calculate it, and a mark that fails the ratio doesn't ship; the system picks a legible variant instead. Any claim that can be computed should never be generated.

We put world-class references in front of the judge. Not style instructions, but actual exemplars, chosen to be adjacent to the work at hand. The scoring got harsher and the output got better, which is the correct direction for those two things to move together.

Every number traces to a source. A generated financial model can produce a confident, wrong figure and wrap it in fluent prose. Our rule is that if a number appears in a deliverable, it points at the upstream work that produced it, and a validator fails the document if it doesn't. Fluency is not evidence.

We added subtraction as a step. Not "make it concise" buried in a prompt, but a separate pass whose only job is to remove, with the previous version retained so nothing good is lost by accident.

The honest part: this is not solved

We'd be doing the thing we're criticizing if we ended on a clean win.

Reference-conditioned judging makes output meaningfully less generic. It does not give a system taste. We know this because our own judge and our own founder disagreed about a logo. The judge preferred the polished, conventional option; the founder preferred the one with a point of view, and the founder was right. We changed the judge.

That's the standing order internally: when the rubric and the human eye disagree, the eye wins and the rubric gets fixed. Any vendor who tells you their evaluation loop has replaced human judgment is describing a system that has stopped improving.

There's also a limit we can't engineer around. A system working from a blank prompt has less to work with than a system working from your actual materials: your existing brand, your real customers, your rejected attempts. Ground truth beats a good guess every time, which is why the most useful thing you can give an AI tool is not a better prompt but better inputs.

Three questions to ask any AI tool you're evaluating

"What is this graded against?" If the answer is "the model evaluates it," you'll get competent-and-forgettable, and the loop will approve it. If the answer names references or standards, ask to see them.

"Which claims are computed and which are generated?" Accessibility ratios, financial figures, market sizes, and citation-backed facts should be computed or sourced. If the same process that writes the prose also produces the numbers, treat every number as a draft.

"What does it remove?" If a tool only ever adds sections, you're getting length, not quality. Ask what its output looks like after editing, and whether the tool does any of that editing itself.

Why this matters more than model choice

Founders ask us which model we use, as though that were the differentiating decision. It isn't, and the models keep changing anyway. Two systems built on the identical model produce output a customer can tell apart, because the difference lives in what surrounds the model: what it's shown, what it's graded against, what it's forbidden to invent, and what it's required to cut.

That's the part that's engineered rather than purchased. And it's the reason "competent" is now table stakes rather than an achievement. Everyone has competent. The remaining question is whether anyone remembers your work afterward.

Common questions on this topic

Why does AI-generated design look so similar across different tools?
Because generation without a strong constraint returns the statistical center of its training data, and a request like 'clean and modern' is a request for exactly that center. Different tools built on similar models converge on the same middle. Distinctiveness comes from constraints and reference standards that pull output away from the average, not from the model itself.
What is reference-conditioned evaluation?
It means scoring generated work against concrete world-class examples placed alongside it, rather than asking a model to rate quality in the abstract. A model grading in a vacuum applies the same average-seeking priors that produced the work, so it approves mediocre output. Shown genuine exemplars and asked to justify why the candidate falls short, the same model becomes a far more useful critic.
Can AI evaluation replace human design judgment?
No. Automated evaluation catches objective failures reliably: contrast that fails accessibility thresholds, unsourced numbers, missing required elements, and it raises the floor considerably. It does not supply taste. When an evaluation rubric and an experienced human eye disagree, the correct response is to trust the eye and fix the rubric.
How can I tell whether an AI tool's numbers are trustworthy?
Ask whether each figure is computed or generated. Accessibility ratios, financial projections, and market sizes should trace to a calculation or a citable source, with validation that fails the document when they don't. If the same generative pass that writes the narrative also produces the figures, treat every figure as an unverified draft regardless of how confidently it is presented.