You have a PMF score. That is the easy part.
Most teams run the Sean Ellis survey, read the percentage, feel briefly good or briefly bad about it, and go back to the roadmap they already had. The percentage is the least useful thing that survey produces.
The survey itself is one question. "How would you feel if you could no longer use this product?" Very disappointed, somewhat disappointed, not disappointed. Count the first bucket, divide by the total, and 40% is the line Sean Ellis found across roughly 100 startups. If you want the full breakdown of the threshold and why product-market fit is a hypothesis rather than a feeling, that post covers the setup.
This one is about the day after.
Because here is what I keep seeing. A team runs the survey, gets 31%, and concludes they do not have fit. Correct, and completely unactionable. Another team gets 44%, declares fit, and changes nothing about how it prioritises. Also correct, also unactionable. In both cases the number was treated as a verdict, and the three hundred free-text answers sitting underneath it were never opened.
That is backwards. The score is a thermometer. The segments are the diagnosis.
What Superhuman actually did
The best documented run of this is Rahul Vohra's, written up in How Superhuman Built an Engine to Find Product-Market Fit. Superhuman started at 22%, well under the line, and reached 58% over roughly three quarters.
The part people quote is the jump. The part that matters is that the jump came from a process they could repeat, and the process is mostly segmentation.
Segment to find your supporters. Take only the "very disappointed" answers and read what those people say about themselves. Not what they want built. Who they are, what job they were doing, what benefit they named. That group is your high-expectation customer, and it is almost always narrower than the market you wrote on your deck.
Analyse in two directions. From the supporters, learn what to protect. From the fence-sitters, learn what is holding them back. From the "not disappointed" group, learn nothing. Vohra's advice there is blunt and correct: politely disregard them. They are not a segment you are failing, they are a segment you are not for.
Filter the middle. This is the step that gets skipped. Not every "somewhat disappointed" user is convertible. Superhuman split theirs by whether the core benefit, speed, actually resonated. The ones who cared about speed were worth building for. The ones who did not were the "not disappointed" group with better manners.
Split the roadmap. Half toward deepening what supporters already love. Half toward removing what blocks the convertible middle. Then run the survey again and watch the number.
What it looked like on our side
Last October we ran the same survey at UX Pilot, the AI design tool I work on, sent to paid users. We came back with 49% across 303 responses. Above the line, which was good news for about ten minutes.
The segmentation is where it stopped being a vanity number:
- 149 users said they would be very disappointed. That is the core market.
- 85 liked the product but named a specific blocker.
- 69 did not fit the value proposition at all.
Reading the 149 gave us an ICP we could actually write down: designers and product managers who need to get from an idea to something visual, fast. Their number one benefit was speed, quick ideation and prototyping. Their number one pain was manual design work eating their week. None of that was a surprise in the abstract. Having 149 people say it in their own words, unprompted, is a different kind of artifact to take into a prioritisation meeting.
The 85 in the middle were more useful still, because they told us exactly what to fix, and they agreed with each other: Figma integration, more control over the AI output, better editing. Those three went onto the roadmap as priorities, and we shipped them.
The 69 we left alone. That is the discipline the whole thing rests on, and it is genuinely hard. A team that treats every unhappy respondent as a gap to close will build a product that is slightly acceptable to everyone and essential to nobody.
Two ways to misread your own results
Surveying the wrong people. The question only means something if the respondent has actually used the product enough to miss it. Send it to your whole signup list and you will measure how many people remember installing something. Send it to engaged, recent, ideally paying users, and you will measure fit.
Treating a low score as a product problem. Sometimes it is. Sometimes the product is fine and the audience is wrong, which produces exactly the same number and a completely different fix. If your "not disappointed" bucket is large and your "very disappointed" users all look like each other, that is not a signal to build more. It is a signal to iterate on who you serve rather than what you build.
If you are below the line
A score under 40% is not a reason to stop. Superhuman started at 22%. It is a reason to go back a step: the survey measures product-market fit, and the checkpoint before it is whether a specific group has a problem painful enough to use your specific solution. That has its own sequence, and running Problem-Solution Fit in five steps is a faster route out of the twenties than another quarter of features.
What you should not do is run the survey once, file the number, and call it measurement. The whole point of Vohra's engine is the last step: repeat it. One score is a snapshot of a market you were guessing at. Four scores in a row is a feedback loop, and it is the only one I know of that tells you whether the last quarter of work moved fit or just moved the backlog.
Turning that loop into an actual operating rhythm, rather than a survey somebody ran once, is the spine of our PMF Program.
So if you already have a score, do not send it to the board yet. Open the free-text answers, sort them into three piles, and write one sentence describing the people in the first pile. That sentence is worth more than the percentage.