Cover image for Taste has to become testable
Alexander Hipp

Taste has to become testable

Defining good AI output is still closer to art than science. I learned this at Miro the hard way. That is the reason to write it down, not the excuse to keep it in your head.

At Miro, we did a couple of experiments with different prompts to Miro AI, several models and MCP, and their attempts at generating a whole visual workspace.

Different people we showed the result to had different favourites. Nobody really had the same. Worse than that, people preferred different parts of different outputs: the structure from one, the wording from another, the layout from a third.

There was no yes and no no. There was individual taste, several people's, pulling in different directions.

For most of my career that got settled in a room, because software behaved the same way twice and the argument was about what to build, not whether the thing in front of us was any good quality.

Nobody in the room is paid to say no

When there is no single verdict to reach, a review stops being a review and becomes a bit more like a social event.

Look at who carries what. The person who says this is not good enough has to hold up a launch (that was me in the beginning :D), argue against momentum, and do it while pointing at a feeling. The person who says it looks fine carries nothing. The most senior reaction sets the tone, and the first opinion spoken anchors everyone else.

That is not a taste problem. It is an incentive problem, and incentives produce the same outcome every time.

If a testing session ends with a decision and no written criteria, nobody reviewed quality. We would just take a vote. This is clearly not scalable.

Then the next feature arrives and the reasoning is gone. Nobody recorded what they were looking for, so nobody can be held to it, including me.

Start from the failures, then build the bar back up

The instinct is to describe excellence. Set a bar. This was also me. Helpful and on brand. Words like that feel like standards and behave like mood, which is how five people agree on a sentence and disagree on every output it was meant to cover.

Start at the other end. Name what the output must never do, specifically enough that everyone recognises the case on sight. Failure space is smaller than excellence, easier to list, and much harder to argue with.

When a room cannot agree on what great looks like, the floor is the only thing everyone can sign.

But this is also not sufficient. A product judged only by what it avoids drifts toward the inoffensive, so the bar for good gets built on top of the floor once the floor holds.

A stack of real outputs and great outcomes marked as good or bad was helping us as a team more than a page of principles, and it drags the disagreements into the open while they are still cheap.

Who marks them matters more than people expect. Domain experts should be the right answer here but they are rarely available, so in practice it is the team, product managers included. You encode your own assumptions about the work and then call them a standard.

What I would love to have in place before anyone builds:

  • the failures we refuse to ship, named
  • the bar for good, written and agreed
  • real outputs, labelled by whoever knows the work best
  • a standing look at what users actually got
  • a read of how teams further ahead do this

The last one gets skipped out of pride. Measuring output quality has real practitioners now, and inventing your own vocabulary buys a worse version of something that already exists.

Criteria are a ledger, not a gate

Launch week arrives, the evals of one of the outputs trips, and everyone looks at you.

In the ideal world.

I would not slip every launch over it, and I will not pretend the criteria work as a hard gate. Severity and reversibility decide it, not the deadline and not who is in the room. Something serious means rework, and the honest reading in that moment is that we learned late, often about the criterion itself. It was too broad, or aimed at the wrong risk, or written by people who had not yet watched the failure happen. A softer miss ships, and we watch production.

The criteria do not make the call. They decide which conversation you are having, and they put whoever waves it through on the record.

Which is also why this cannot sit with one role. Product knows what good means to the person on the other end. Engineering knows what can be measured, what it costs, and how the measurement will lie. One document, both names on it, neither side allowed to hand the problem back.

The art stays at the edges

Most of what makes an output good can be written down. The edges cannot, and I would rather say that than oversell the method.

For open-ended visual output, those edges are not rare. A wall of outputs where everyone prefers a different fragment is the normal case, and no score resolves it. You can measure whether a layout is legible.

You cannot measure whether this particular arrangement is the one a team will think with.

There is a quieter failure too. Every number improves while the product gets worse, because the criteria have stopped describing the work and started describing themselves. The only defence is someone who still reads raw output for no measurable reason.

So the scores are not the point. They buy back attention. They cover the part of quality that repeats, which leaves your taste for the part that does not.

Write down what you can be wrong about

Judgment was always the job. What changed is that it can no longer live only in my head, because the thing I am judging does not behave the same way twice, and neither do the people judging it with me.

Written criteria are a bet I can lose. Taste I never write down can never be wrong.

A standard you can be proved wrong about is the only kind worth having.

Just a couple of thoughts on taste and AI evaluation.

Article by Alexander Hipp (AI Product Lead)
Read more about AI, Evals, Product Management