If it’s possible to train an AI model to have good taste in design, it may be time to rethink our training techniques and approaches. All of our efforts to train models for higher competence in the task of building websites may be discouraging the model behaviors most aligned to human preferences when evaluating the final results.
To explore the current landscape of AI design, we ran a study: 10 models from 10 open and frontier labs were each provided 10 "client" briefs and asked to record their creative decisions and supporting rationale as they selected brand fonts, colors, imagery, and design patterns. Then we had human annotators evaluate and compare all the outputs.
The findings were pretty fascinating. In a series of articles this month, we’ll explore the questions around where the models converge and diverge, how familiarity and novelty influence human preferences, and if AI taste be bought with more tokens and higher model effort level.
Any of today’s models can build responsive, working sites. Most AI design benchmarks test whether they can reproduce something that already exists: give a model a reference interface and see how closely it can rebuild it.
But that exercise doesn’t tell us much at all about taste.
In this study, there were no reference designs. Each model had to make its own decisions about color, type, and layout, without the ability to use outside tools, frameworks, or packages. We also asked each model to explain those decisions. Then we showed the results to people and asked which ones they preferred.
As frontier models exhibit rapidly increasing levels of competence and skill in execution, how do we train for the more nuanced and unstable domain of human taste and preference?
As Michael recently put it in his writing on taste, the advantage increasingly shifts from simply having access to capable models to knowing what good looks like.
That’s a much harder thing to evaluate. There’s no compiler for taste.
We gave all the models the same 10 creative briefs, ranging from a telemedicine app to a New York skatewear brand.
We used the Flue framework to give every model the same controlled agent environment and OpenRouter for a common interface, pinning each request to a named provider with fallbacks disabled. We controlled tool usage, forbid web search (often to angry complaints from the models), and forbid the use of frameworks, component libraries, etc.
For each brief, every model went through three arms:
All 10 models used the same image-generation model, so differences in image quality wouldn’t explain the results.
Finally we asked people to judge the results. To do the human evaluation, we ran a blind, balanced round-robin in Label Studio. Every pair of models went head-to-head on every brief, giving us 450 matchups and 4,500 votes from 29 annotators.
Annotators saw two sites side by side. They didn’t know which model made either one. They picked a winner or called it a tie. The goal was to separate “this works” from “I prefer this.”
Across 100 generated sites, three patterns show up again and again. The common elements won’t surprise anyone who’s familiar with the tells of AI-generated web design:
But brand color selection showed the most surprising convergence pattern.
Example: For the telemedicine brief, the models scored 0.95 on primary-hue concentration, where 1.0 means every model chose the same hue. Eight of ten models chose a deep teal as the primary color, and several explained it as a way to stand apart from "generic medical blue". The models converged while each believed it was diverging.
Independent film was the opposite. It scored 0.14, making it the most varied brief in the study.
When models did disagree, they tended to split into a couple of predictable camps.
Ten different labs, and we keep ending up with the same two or three answers.
The most interesting part is that this convergence doesn’t begin with CSS.
In the direction arm, there was no code, framework, or component library. The models were simply describing what they wanted to make. Yet the convergence was already there.
Two of the 10 models emitted Tailwind color tokens. The other eight didn’t. But all 10 still closely orbited around the same hues. In short, it’s not a tooling issue: the defaults chosen by Tailwind are embedded in the training data and influence the outputs, even when Tailwind isn’t part of the toolchain.
Their reasoning looked similar, too.
Of the 563 color rationales, 36% described a choice in terms of what the model wanted to avoid: “without X,” “rather than Y,” and similar constructions.
The models really were trying to differentiate their designs. Yet they kept arriving at eerily similar ideas of what “different” should look like.
If every model has access to roughly the same capabilities, differentiation has to come from somewhere else. And if the models’ default judgments converge, then capability alone isn’t going to produce distinctive work.
We’ll dig more into those rationales—and the references behind them—in a later post.
If your product generates design, this house style is likely part of the default. And a traditional competence benchmark probably won’t tell you that.
AI "slop" used to be rooted in competence or execution problems. Now, AI slop is what happens when models fail to align with human taste and keep pace as that taste evolves.
If taste is becoming a moat, then taste needs an evaluation loop. Show people the outputs. Ask them to choose. Measure where they agree, where they disagree, and what model behaviors consistently lead to choices that people prefer.
That’s what we did here, and we’re eager to talk to other teams researching the intersection of AI and human preferences.
Next, we’ll publish what our annotators chose across every blind matchup, which models came out ahead, and which behaviors were rewarded with human preference. Stay tuned.