The Sean Ellis Test: Measuring Product-Market Fit With One Question
Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
Most definitions of product-market fit are vivid and hard to measure. "You'll know it when you see it" is a nice sentiment and a poor operating metric. The Sean Ellis test is the opposite. It's a single survey question, a three-way answer scale, and one number to compare against a benchmark.
The question asks users how they'd feel if they could no longer use your product. The share who answer "very disappointed" becomes your score. Ellis, an early growth leader at companies like Dropbox, LogMeIn and Eventbrite, argued that a score of 40% or more is the line that separates products worth scaling from products still hunting for a reason to exist.
This article covers the exact wording, where the 40% line came from, who you should survey, how Superhuman turned the score into a working method, and the limits you need to respect. If you want the broader idea first, read product-market fit explained.
The question and the answer scale
The survey wording, as Superhuman used it and as Ellis described it, is:
How would you feel if you could no longer use [product]?
- Very disappointed
- Somewhat disappointed
- Not disappointed
That's the core of it. The test deliberately doesn't ask whether people are satisfied, whether they'd recommend you, or whether they like the design. Satisfaction is cheap. Dependence is expensive. A user can be pleased with your product and still shrug if it vanished tomorrow, and the "very disappointed" answer is meant to isolate the people who would genuinely miss it.
Ellis's own blog states the threshold plainly. In a 2010 post on optimization, Startup Marketing, he recommended that startups not begin optimizing "until at least 40% of your randomly surveyed users say they would be 'very disappointed' without your product." In an earlier 2009 post on when a startup should start charging, he defined product/market fit as at least 40% of your active users saying they would be "very disappointed" if they could no longer use the product.
Two details in those posts are easy to miss. First, Ellis tied the score to a decision: don't pour effort into funnel tuning until the core product clears the line. Second, he wasn't claiming a law of nature. He was describing a pattern he'd seen across the companies he worked with.
Where the 40% number came from
The benchmark is empirical, not derived from theory. In a reader comment on his charging post, Ellis said he couldn't share the list of companies he'd compared. First Round Review's write-up of the method says Ellis arrived at 40% after benchmarking nearly a hundred startups, finding that companies struggling for growth almost always had fewer than 40% of users answer "very disappointed," while companies with strong traction almost always exceeded it.
Notice the word "almost." The threshold is a rule of thumb drawn from one practitioner's sample, and the sample isn't public. That's fine for a decision aid. It's a problem if you treat 39% as failure and 41% as success. Think of 40% as a zone where the odds change, not a switch.
Key Facts
- The question: "How would you feel if you could no longer use [product]?" with three answers: very, somewhat, not disappointed.
- The score is the percentage answering "very disappointed."
- Ellis's published rule of thumb is at least 40% "very disappointed" among surveyed users.
- Ellis says he generally looks for at least 30 responses, where results start to stabilize.
- Superhuman reported a starting score of 22%, which rose to 58% after three quarters of work.
- The score is self-reported, a snapshot, and sensitive to who you survey.
Who to survey
The score is only as good as the people answering. The wrong sample can make a weak product look strong or the reverse.
Ellis's recommendation, as relayed by Superhuman's founder, was to survey users who've recently experienced the core of the product. Superhuman applied that by polling people who'd used it at least twice in the last two weeks. That filter matters for two reasons.
- Sign-ups aren't users. Someone who registered during a press spike and never came back can't tell you what they'd miss. Their answer is noise.
- Old users aren't current users. A person who used the product heavily last year and stopped is a different data point from someone using it this week.
On sample size, Ellis answered a reader directly in the comments of his charging post. He said he generally looks for at least 30 responses, where results start to stabilize, and that more is better. Superhuman's founder said you start to see directionally correct results at around 40 respondents. Those are practitioner rules of thumb, not statistical guarantees. We'll put numbers on the uncertainty in the worked example below.
Ellis's wording also says "randomly surveyed." If you only email your most enthusiastic customers, you'll inflate the score and learn nothing. If you can, sample across the recent-active user base without cherry-picking.
The Superhuman case
The best-known use of the test is Superhuman's, described by founder Rahul Vohra in a First Round Review article. It's worth reading because it shows what to do after you get a bad number.
Superhuman emailed recent active users a short survey with four questions:
- How would you feel if you could no longer use Superhuman? (very, somewhat, not disappointed)
- What type of people do you think would most benefit from Superhuman?
- What is the main benefit you receive from Superhuman?
- How can we improve Superhuman for you?
The result was 22% "very disappointed," which the article says made it clear the company hadn't reached product-market fit. The numbers the article reports afterward:
| Stage | Score ("very disappointed") |
|---|---|
| Starting point, summer 2017 | 22% |
| After narrowing to the personas who loved the product | 33% |
| After three quarters of product work | 58% |
The segmentation method
The jump from 22% to 33% didn't come from changing the product. It came from changing the question being asked of the data. The article describes the steps this way:
- Group by answer. Split responses into very, somewhat and not disappointed.
- Assign a persona to each respondent. Then look at which personas show up in the "very disappointed" group.
- Narrow the market to those personas. Re-scoring only the segments that loved the product raised the score to 33%, because responses from personas outside the target market no longer counted.
The article also reports that the team used answers to question two, from the very disappointed group only, to describe their most discerning ideal customer in the users' own words.
Fixing the "somewhat disappointed" group
The second half of the method is what makes it a working engine rather than a one-off measurement. The article describes three moves:
- Study the fans. Read what the very disappointed group says is the main benefit. At Superhuman, they named speed, focus and keyboard shortcuts.
- Ignore the "not disappointed" group. The article argues their feedback leads to distracting feature requests and a muddled roadmap.
- Split the "somewhat disappointed" group. Superhuman separated people for whom speed was the main benefit from those for whom it wasn't, and used the first group's complaints to decide what to build to move them up a notch.
The takeaway isn't "copy Superhuman's roadmap." It's that the test tells you where to look. The people who love you show you what's working. The people on the fence show you what's blocking growth. Everyone else is a distraction at this stage.
How to run it
Here's a practical sequence, assembled from the Ellis and Superhuman sources above.
- Pick the audience. Active users who've used the core feature recently, such as twice in the past two weeks. Don't include people who only signed up.
- Send the question. Use the exact wording and the three-option scale. Changing "disappointed" to "unhappy" or adding a fourth option makes your score incomparable to the benchmark.
- Add follow-ups. Ask what type of person would benefit most, what the main benefit is, and how to improve. These are the questions Superhuman used.
- Collect enough responses. Ellis's stated floor is around 30. More is better.
- Score it. Divide "very disappointed" by total responses.
- Segment before you conclude. Break the score down by persona, plan type or use case. A blended 28% may hide a 55% pocket.
- Repeat. Re-run the survey after meaningful product changes. One reading tells you where you are, a series tells you whether you're moving.
A hypothetical worked example
This is an invented illustration, not a real company. Suppose a project-planning tool surveys 60 recently active users and gets these answers:
| Answer | Count | Share |
|---|---|---|
| Very disappointed | 21 | 21 / 60 = 35% |
| Somewhat disappointed | 24 | 24 / 60 = 40% |
| Not disappointed | 15 | 15 / 60 = 25% |
The score is 35%, below the 40% line. To reach 40% on the same 60 responses, the team would need 0.40 x 60 = 24 "very disappointed" answers, so three more people would need to move up.
Now segment. Suppose the 60 respondents split into two personas:
| Persona | Respondents | Very disappointed | Score |
|---|---|---|---|
| Agency project managers | 25 | 14 | 14 / 25 = 56% |
| Freelancers | 35 | 7 | 7 / 35 = 20% |
The blended 35% hides two very different stories: 14 + 7 = 21 across 25 + 35 = 60. The agency segment is well above the line and the freelancer segment is far below it. The sensible reading is to focus on agencies first, learn why they'd miss the tool, and stop designing for freelancers for now.
One more caution on the arithmetic. With 60 responses and a true score near 35%, simple binomial math gives a 95% margin of error of roughly 12 percentage points (1.96 x the square root of 0.35 x 0.65 / 60). That's our own calculation for illustration, not a published Ellis figure. It means a reading of 35% is statistically compatible with anything from the low 20s to the high 40s, which is why a single small survey shouldn't settle an argument.
Limits of the test
The test is popular because it's simple, and its simplicity is also its weakness.
It measures stated feelings, not behavior. People say they'd be very disappointed and then churn anyway. Or they say they'd be fine and use the product daily. Self-reported intent is a leading indicator at best.
It's sensitive to who answers. Survey only power users and you'll inflate the score. Survey everyone who ever signed up and you'll deflate it. Response bias also matters: the people who bother to answer a survey tend to be more engaged than the people who don't.
It's a snapshot. One score says where you stand today. It doesn't show whether you're improving, and it can't tell you why. Superhuman's article makes the point that the score is something to keep tracking as you grow, because new kinds of users arrive and early adopters are more forgiving than later ones.
The 40% line is a rule of thumb. It comes from one practitioner's unpublished sample. A product at 38% isn't doomed and a product at 42% isn't safe.
Price and context matter. Ellis himself noted in a comment that he starts from a price of zero, because if people aren't disappointed to lose a free product, they'd be even less bothered once it costs money. A high score on a free product still has to survive a real price.
The practical fix is to pair the survey with behavioral data: retention curves, repeat usage, and whether users come back unprompted. If 45% say "very disappointed" and your cohorts still fall off a cliff in week three, believe the cohorts. For how retention fits into the bigger picture, see product-market fit for SaaS, and for choosing the one number your team rallies around, see the North Star metric.
When the score is low
A low score isn't a verdict. It's a prompt to learn. Three common responses:
- Narrow the market. If one segment scores high, double down there, as Superhuman did.
- Change the product. If nobody names a clear main benefit, the value proposition may be fuzzy. That's the point where teams consider a pivot.
- Go back to the customer. If you can't explain who loves you or why, you probably skipped some conversations. Customer discovery is the fix.
Remember too that the test works best on products with actual users. Before you have any, you can't survey your way to fit. You have to build something people can use first.
Frequently Asked Questions about the Sean Ellis Test
What is the Sean Ellis test?
It's a one-question survey that asks users how they'd feel if they could no longer use your product, with the answers very disappointed, somewhat disappointed, or not disappointed. The percentage who say "very disappointed" is your score. Sean Ellis recommended 40% as the threshold that suggests the product is a "must have."
Where does the 40% benchmark come from?
Ellis described it on his Startup Marketing blog as at least 40% of surveyed users saying they'd be very disappointed. First Round Review reports he set the figure after benchmarking nearly a hundred startups, though the list isn't public, so treat it as a practitioner's rule of thumb rather than a proven law.
How many people do I need to survey?
Ellis said in a blog comment that he generally looks for at least 30 responses, where results start to stabilize. Superhuman's founder said you get directionally correct results around 40 respondents. Small samples carry wide margins of error, so more is better.
Who should I send the survey to?
Send it to people who've recently used the core of your product, not everyone who ever signed up. Superhuman surveyed users who'd used it at least twice in the previous two weeks. Avoid cherry-picking your happiest customers, because that inflates the score.
What did Superhuman's scores look like?
According to founder Rahul Vohra's First Round Review article, Superhuman started at 22% "very disappointed" in summer 2017. After segmenting to the group that loved the product it reached 33%, and after three quarters of product work it reached 58%.
What are the main limits of the test?
It captures self-reported intent, not actual behavior, and it depends heavily on who you survey. It's also a snapshot in time. Pair it with retention and usage data before you make big bets on the result.
Related reading
