Contact Sales
Skip to essay

EssaysAI evaluation

Is taste a moat?

A framework that shows when evaluation serves as a quality floor and when it functions as a defensible learning system.

By Michael Malyuk · CEO, HumanSignal · 12 min read

I’ve been talking with other founders about what counts as a moat these days, and many of us keep coming back to the word “taste”. Once everyone has the same models, the only thing left to compete on is knowing what good looks like.

On the surface, it seems right. The argument is that frontier models are trained with reinforcement learning against verifiable rewards (RLVR). To teach these models something, you need a way to check whether it did it correctly, hence the “verifiable” and “reward” in the name. For example, code produced by the model either compiles or not, the test passes or fails. But taste has no compiler, so it’s unclear how to train for it. It stays scarce and can be claimed as a moat.

A few people have thought about this deeply, pointing out that if taste is pattern recognition over accumulated exposure, models will eventually learn it because pattern recognition is what models excel at [1]. Another argument is that taste isn’t a moat but more of an alpha, an edge that decays with time and is valuable only relative to the baseline that is always rising [2].

While all these perspectives are valuable, I started thinking about another question: can you turn taste into a system that keeps up to date with the world so you always have alpha? To answer that, I want to look at taste and moat from two angles: first, how to extract the signal, and second, the timelines for when that signal is valid and how long it stays “uncopyable”, thus creating a moat.

Can taste be learned?

There are many ways to turn taste and judgment into a signal the model can learn: pairwise comparisons that train reward models, rubrics that drive LLM judges, direct ratings from experts. The mechanics may differ, but they all produce the same thing: a number (or set of numbers[3]) attached to an output indicating how good it is, in other words, a score.

Once you have that signal, you can do one of two things: use it to understand the system or use it to improve the system. Used diagnostically, the signal becomes an eval. You run a set of tasks, examine the scores, and use them to understand how the system behaves. Used as an objective, the signal becomes a target. You compare experiments, decide which model to ship, or use techniques such as RLHF to reinforce the behaviors that make the score go up.

A benchmark may begin as a way to measure behavior, but once it influences which experiment wins or which model ships, people start optimizing against it. The measurement begins to shape the system it was designed to evaluate.

That leads to a more important question: After the system learns to satisfy the signal, does that signal still represent the quality you originally cared about?

Three problems tend to get in the way:

  • First, repeatability does not mean validity. A judge might give the same score each time but still measure the wrong thing. A coding agent may pass the tests even though it leaves a mess. A support agent can receive high ratings for being helpful even if the tickets end up being reopened. In such cases, where the outcome you care about only becomes apparent after a delay (weeks later), the loop takes weeks to complete, which means it takes a long time to make adjustments and fix the problem. So consistency matters, but validity matters at least as much.
  • The second is Goodhart’s law[4], which I alluded to earlier. When models are optimized based on the signals you give them, rewarding a coding agent whenever tests pass might lead it to produce weak tests rather than improve the code. Likewise, if you introduce a style judge, the agent will pick up on the judge’s tastes rather than yours, provided that the judge wasn’t fully aligned with you. Hence, when designing evaluation systems, it is important that they be able to withstand optimization.
  • The third is that consensus isn’t yet a reward function. Preference data mostly captures what the majority agrees is fine, typically within well-understood dimensions like clarity and helpfulness. That creates a floor but not an edge. Averaging multiple opinions removes obvious failures but may smooth out valuable differences and unusual points of view. The valuable signal is usually contextual and multi-dimensional: good for whom, in which workflow, under what constraints?

How defensible is it?

Now suppose your signal is honest: your eval isn’t lying, and you can learn from it. Does that signal alone give you a moat? Not yet. You also have to keep timelines in mind, since signals tend to decay, expire, and become less valid. Three timelines matter:

  • The learning timeline says feedback time plus update time must be shorter than the time it takes the target to move. In most practical situations, this is hard to measure, even though it looks measurable on the surface. You typically learn about drift only after labels stop working in production.
  • The imitation timeline says you have to update your edge faster than rivals can copy you, for example, by buying the same dataset or distilling your models. If you ship a model improvement, how long does it take for your competitor to match it? If you do not own the signal you’ve learned from, it’s a commodity. If it’s owned but not updated, it tends to become a depreciating asset.
  • The absorption timeline is new and was not part of pre-LLM-era moat discussions. This is where your edge can disappear without competitors copying it. The base model simply improves. A custom calibration you did can now ship as default behavior. Your moat must compete with both: competitors trying to copy it and the baseline absorbing it.

Mapping the Learning Loops

We can combine these timelines into a useful map. The horizontal axis is the learning timeline: can trustworthy feedback improve the system before the target changes? The vertical axis combines the imitation and absorption timelines: can the advantage be renewed before competitors copy it or the base model absorbs it?

Learning ability increases left to right; defensibility increases bottom to top. Four quadrants: Story, Learnable Commodity, Structural Moat, and Learning Loop. Production feedback pushes toward the Learning Loop; Goodharting moves left; absorption moves down.
View the map at full size ↗

This creates four kinds of positions. Keep in mind that an organization can run multiple loops at once, and each loop moves over time: Goodhart’s Law drags loops to the left, the rising base model drags them down, and proprietary feedback pushes them up and to the right.

  1. Story: weak learning, weak defensibility. This is “taste” that lives mainly in the founder’s head, an advantage that exists as narrative rather than as a system. There may be good judgment involved, but the organization cannot reliably distinguish skill from luck, reproduce the judgment, or carry it forward. The advantage has to be recreated each time. I call it a “Story” because it’s usually what is left after.

  2. Learnable Commodity: strong learning, weak defensibility. Everyone can learn from the signal, so everyone rises together. This is the quality floor: necessary, but not an edge. SWE-bench illustrates the lifecycle. When the original benchmark launched in October 2023, the best models resolved less than 2% of its issues. SWE-bench Verified, a human-validated subset released in 2024, later became a standard measure for coding agents, but OpenAI eventually stopped reporting it after finding both flawed tests and evidence that frontier models had been exposed to benchmark problems or solutions during training.[5]

    That doesn’t mean the benchmark was useless. It means its success changed its function. It helped raise the baseline while gradually losing its ability to measure the frontier.

    This is where companies often confuse expense with defensibility. An eval may be costly to build without becoming a moat. If competitors can observe the standard and inherit the lesson without paying the same cost, the original expense does not protect you.

  3. Structural moat: weak learning, strong defensibility. This is an advantage others cannot easily access, but it doesn’t come from learning. Distribution, switching costs, brand, etc: real moats, just not “taste originated” moats. These moats can last pretty long, and defend a position, but without making the product any better.

  4. Learning Loop: strong learning, strong defensibility. This is the most interesting quadrant. The system receives meaningful feedback from production, converts failures into new evaluations or training data, and ships improvements while those lessons are still valuable. Because the feedback comes from an owned workflow, competitors cannot easily reproduce it. The moat is not the original judgment or dataset. It is the system’s ability to keep producing the next one.

Copilot, Cursor, and what makes a loop defensible

The third and fourth quadrants are best illustrated by running a thought experiment comparing Copilot vs Cursor.

For years, Copilot appeared to have the ingredients of a formidable data advantage: broad distribution, extensive IDE usage, and feedback from developer interactions. Yet Cursor built a competitive product without beginning with the same history of proprietary usage.

That doesn’t mean Copilot had no moat. Its distribution through GitHub, VS Code, and Microsoft’s enterprise presence was itself a significant advantage. But the comparison highlights an important distinction: access to feedback is not the same as converting feedback into product improvement.

Cursor has made that conversion explicit. It built CursorBench from real internal coding sessions and combines offline evaluation with controlled experiments on live traffic [6]. It also describes a reinforcement-learning system that can collect production signals, train and evaluate a new Composer checkpoint, and deploy an improvement in roughly five hours [7].

Cursor has not found a perfect reward function. Its own account describes cases in which the model learned unintended behaviors and the team had to revise the system. That ability to detect failures and respond quickly is the point.

In its simplest form, the operating loop looks like this:

Output → judgment → intervention → deployment → observed outcome → revised judgment

The intervention might be a training update, a different model, a new prompt, a tool-policy change, or another product decision. The difficult part is not drawing the loop. It is ensuring that the judgment reflects real quality, the intervention happens quickly, and the outcome makes its way back into the next decision.

That is what can make the loop defensible: not simply having feedback, but repeatedly converting proprietary feedback into a better product.

Who needs to own the loop?

A closer look at that map suggests the moat belongs to the company running production: collecting and storing traces, plus owning the workflow. And indeed, for an AI-native company, the loop is the product. The harness and the training pipeline are the business, and hiring researchers to run them is the point. But for everyone else, for example, a bank deploying a screening agent, or an insurer triaging claims, the loop is a quality and governance function that wraps around somebody else’s model.

That second case is where it gets interesting, because the model vendor now sits on your absorption timeline. Any calibration generic enough to help everyone will eventually show up in their next release, whether or not they ever see yours. So what does the bank actually need to own? Not the model. It needs to own the parts that encode its own judgment and can’t ship in anyone else’s release: traces from its workflows, rubrics and judges calibrated to its risk tolerance, a taxonomy of failures its reviewers have adjudicated, and a record of which of those judgments turned out to predict real outcomes. The model can be rented. The judgment about whether it’s working shouldn’t be.

Taste still matters, and matters a lot, but one level higher than where the conversation usually puts it. Taste decides which failures deserve a category, which disagreement contains information, which proxy is safe to optimize, or when a benchmark has been quietly saturated. But for meta judgment to be a moat, it has to leave behind something that compounds. Otherwise it’s an individual intuition applied fresh each time, which is alpha, which is decaying and waiting for the next model release.

If you’re building one of these loops, the directions on the map translate into work. For example, to move right, improve the connection between your eval and reality: narrow task distribution, validate fast proxies against the slow truths they claim to predict, and collect rejected actions counterfactually. For a support agent, that means tying rubric scores to resolution, escalation, repeat contact, and eventually retention. For a coding agent, that means tying test success to review acceptance, reverts, and incidents. You don’t have to wait for the distant outcome: build a ladder of fast signals, keep calibrating it against the slow ones, and the feedback cycle gets shorter.

To move up, collect feedback from your own production experience and build a record of private failure modes, along with a taxonomy of errors and disagreements that have been adjudicated. You might start with a private test set, but the real aim is to keep a system that converts deployment failures into evaluations and then evaluations into shipped changes. This is much more difficult to reproduce.

If this framework is right, it makes a few testable predictions. Expert and trajectory-level judgment gets more expensive while generic output preferences get cheaper, and the spread keeps widening. Products whose loops sit in the quality-floor quadrant converge on the same quality and end up competing on structural moats like distribution and price. Loop owners will show higher deployment frequency and faster regression fixes. And the framework is wrong if generic frontier judges, applied cold, match privately calibrated loops. That would mean absorption has outrun renewal everywhere. I expect that to happen in some domains, but not universally.

So, is taste a moat?

Taste can create an initial advantage. On its own, however, it decays the way alpha does: the baseline rises, competitors adapt, and the edge gets arbitraged away. The defensible asset is the loop that decides what “better” means, checks that judgment against reality, and updates before the signal expires. A simple test for an AI business is: if a competitor can query the same frontier model and observe the same public outputs, what is left? If the answer is nothing, taste was only an input.

Judgment becomes durable through what it builds: the tests, standards, data, and systems that let an organization keep learning and keep generating the next edge after the original insight becomes obvious.

In the end, the moat is what taste leaves behind.

References

  1. When AI has better taste than you
  2. Taste is not a moat
  3. Direct Preference Optimization
  4. Goodhart’s law
  5. Why we no longer evaluate SWE-bench Verified
  6. CursorBench
  7. Real-time RL for Composer