Field Notes  /  Product Discovery
Product Discovery

How strong is your evidence? Score the idea before you argue about it

Two people argue for an hour about whether to build a feature. What they are actually disagreeing about is how much the evidence is worth, and nobody in the room says so. There is a tool for that.

Watch a product prioritisation meeting closely and you will notice something odd. People rarely disagree about what the idea is. They disagree about how much to believe it, and almost nobody says that out loud.

So the argument gets settled by other means. Seniority, volume, whoever spoke last. Everyone leaves the room and the idea proceeds, carrying a confidence level nobody actually measured.

The tool

Itamar Gilad's Confidence Meter fixes that by doing one blunt thing: it ranks types of evidence by strength, so you can point at where your idea sits.

The bands run roughly like this, weakest first.

Opinions. Your own conviction that this is a great idea. Someone else's conviction. A slide in a pitch deck. The fact that it fits a theme leadership likes. This is the weakest evidence there is, and it is where most roadmap items begin.

Assessments. Estimates, projections, business cases, plans. These feel much stronger than opinions because they contain numbers. They are still opinions, with arithmetic applied.

Data. Market data, and evidence from your actual users: what they do in the product, what they said in interviews, what your analytics show about the problem. Now you are working with something outside the building.

Test results. What happened when real users met a real version of the idea. A prototype test, an experiment, a live rollout with a measured outcome. This is the strongest evidence you can hold, and the only kind that survives someone disagreeing with you.

I have used this for years, in day-to-day product work and in the discovery workshops I run. Its value is not really the score. It is that it gives a team shared language for something they were previously arguing about by proxy.

Nobody in a meeting says "my evidence is weaker than yours". They say "I disagree". The Confidence Meter turns the second sentence back into the first.

Why it repairs ICE

Plenty of teams score ideas with ICE: Impact, Confidence, Ease. Impact and Ease get argued about reasonably well, because people have a rough shared sense of what a big feature and a big result look like.

Confidence is where the framework quietly falls over. Everyone scores it by gut, the scores are not comparable between people, and the most confident person on the team is often the one with the least evidence. A number sourced from feelings gets multiplied into the priority order as if it were data.

The Confidence Meter gives that leg something to stand on. Confidence stops meaning "how sure do I feel" and starts meaning "what kind of evidence do I have, and how strong is that kind". Two people can now disagree about a specific, checkable thing.

What it changes with stakeholders

This is the use I did not expect and now rely on most.

When a stakeholder asks why the team is spending three days on prototype tests instead of building, the honest answer is about evidence strength, and without a shared scale that answer sounds like process defence. With one, it is concrete: the idea currently rests on opinions, the test moves it two bands up, and here is what we will know afterwards that we do not know now.

It also makes discovery legible. Interviews, surveys, prototype tests, and fake door tests stop looking like a menu of activities product people enjoy, and start looking like what they are: different ways of buying evidence at different prices. That is the same reason I keep pushing teams to show the work rather than defend the conclusion. The scale is what makes the work legible to someone who was not in the room.

Try it on your own backlog

Take the top three things on your roadmap right now. For each one, write the single strongest piece of evidence behind it, in one sentence. Then put it in a band.

Two things usually happen. Most items land in opinions or assessments, which is uncomfortable but useful. And the exercise takes about fifteen minutes, which makes it hard to argue you did not have time.

Then ask the only question that matters next: what is the cheapest thing we could do this week to move this idea up one band? Usually it is five customer conversations, or a prototype in front of eight people, or a fake door. Small tests, run continuously, are what a culture of experimentation is actually made of, and this is a clean way to decide which test is worth running first.

Two ways to get this wrong

Treating the score as the decision. It is not. Some ideas are worth doing on weak evidence because they are cheap, reversible, and quick to unwind. The score tells you how much risk you are carrying, not whether to carry it.

Demanding test results for everything. If a change takes a day to build and an hour to reverse, the experiment costs more than the mistake. Save the strong evidence for the bets that are expensive to unwind, which is where being wrong actually hurts.

Used well, this sits underneath the same discipline as building on evidence instead of a hunch: know which of your beliefs are load-bearing, and go buy better evidence for those first.

It is also one of the first things we install in a Product Discovery Workshop, because a team that can name the strength of its own evidence argues very differently from one that cannot.

Next time an idea meeting stalls, skip the debate and ask one question. What is the strongest evidence we have for this, and what would it cost to get better evidence by Friday?

Aleksander Uznański
Aleksander Uznański
Founder of ProductTrio. He runs discovery workshops where teams learn to say out loud how much their evidence is actually worth.

Ideas getting picked by whoever argues hardest?

That is fixable, and it is mostly a vocabulary problem. Book a free intro call and bring the argument you are stuck in.

Book a free intro call
Free · 20 minutes · No pitch deck, just your actual problem