Testing With Five Users Is Right. The Way Most Teams Quote It Is Wrong.


person writing on white paper

In 2000, Jakob Nielsen published a short article with one of the most quoted, and most misquoted, claims in product work: “Why You Only Need to Test with 5 Users.” The number stuck. The reasoning behind it mostly didn’t. Twenty-six years later I still hear “five is enough” used to justify decisions Nielsen never argued for, usually in a budget meeting, usually by someone who has never read the original piece.

The gap between what the rule says and how it gets used is worth closing, because the misquote quietly caps how much a team learns before it commits engineering time to a build.

Where the number actually comes from

The five-user figure is not a hunch. It comes from a mathematical model Nielsen built with Thomas Landauer in a 1993 paper, and it looks like this:

N = 1 − (1 − L)^n

N is the proportion of usability problems you will find. The variable n is the number of users you test. L is the average probability that any single user runs into any given problem. Across a large set of projects Nielsen and Landauer studied, L came out to about 31%.

Plug 31% into the formula and test five people, and you land at roughly 85% of the usability problems in the interface. Test one more and the curve flattens hard. That flattening is the whole argument. The sixth, seventh, and eighth users mostly surface problems the first five already showed you, so the money spent on them buys very little new information.

That part is sound. The math holds. The trouble starts with everything the number quietly assumes.

The assumption almost everyone drops

Read the 2000 article closely and you find Nielsen was not telling you to test five people and stop. He was telling you to test five people, fix what you found, and test five more.

His actual recommendation for a fifteen-user budget: run three separate studies of five, not one study of fifteen. His words are blunt about why. “The real goal of usability engineering is to improve the design and not just to document its weaknesses.” One big study documents. Three small rounds let you fix the obvious problems after round one, then watch round two reveal the issues that were hiding underneath them, then use round three to probe deeper structure like information architecture once the surface friction is gone.

So the honest phrasing of the rule is: five users per round, across iterative rounds. The version that reaches most planning meetings is: five users, total, once. Those are not the same claim. The first is a discipline. The second is a shortcut that borrows the first one’s credibility.

In two decades running IT operations and later doing fractional COO work, I sat through a lot of these conversations. The pattern was almost always the same. Someone wanted to cut the research line to protect the timeline, “five users” gave them a respectable-sounding ceiling, and the team shipped on one thin round of testing. When the support tickets came in later, nobody connected them back to the round of research that never happened.

What a study of sixty users showed about betting on five

The strongest counterweight to a lazy reading of the rule is a 2003 study by Laura Faulkner, published in Behavior Research Methods under the title “Beyond the five-user assumption.” Faulkner ran usability tests with 60 participants, then repeatedly pulled random samples of different sizes from that pool to see how much the results swung depending on which people happened to be in your five.

The variance was the finding. With random sets of five users, some sets caught 99% of the known problems. Other sets of five caught only 55%. Same product, same task, same total pool. The only difference was which five people walked in the door.

Fifty-five percent. On a bad draw, half your problems stay invisible, and you have no way of knowing at the time whether you drew well or badly.

Faulkner then showed what buying more users actually purchases: not a higher average, but a higher floor.

  • With 10 users, the worst-case set still found 80% of problems.
  • With 15 users, the average climbed to 97%, and the floor rose to 90%.
  • With 20 users, the worst set found 95%.

Read that as a risk statement, not a coverage statement. Five users gets you a good expected value with a terrible worst case. More users narrows the range of outcomes. Whether that insurance is worth paying for depends entirely on what a missed problem costs you. A missed problem in an onboarding flow is an annoyance. A missed problem in a checkout path, a medical dosing screen, or a financial transfer is a different category of expensive.

The number was never one number

The deeper issue is that “how many users” has no single answer, because it depends on a question most people skip: what are you actually trying to learn?

You’re finding problems in one flow, for one type of user. This is the case the five-user rule was built for. Qualitative, observational, one reasonably uniform audience running one task. Five per round, iterate, and you are on solid ground.

You have genuinely distinct user groups. The formula assumes one population with one shared L. The moment your product serves an administrator and an end user, or a buyer and a spender, or a clinician and a patient, they hit different problems on different paths. Five total does not cover two groups; it covers one group badly and the other by accident. You need a small round for each distinct group, which is why segmenting who you recruit matters as much as how many. The recruiting mistake compounds here: five convenient users from one segment can look like a clean result and be nearly useless, the same trap I described in how survivorship bias eats product discovery.

Your discovery rate is low. L is not fixed at 31%. On a polished product where problems are rare and subtle, the per-user probability of hitting any given issue drops. If L falls to 20%, you need about nine users to reach the same 85%. At 10%, you need roughly eighteen. Mature products have lower L values than early prototypes, which means the more refined your design, the more people it takes to find what is left.

You need a number, not a problem list. This is the line the rule was never meant to cross. If the question is “does version B convert better than version A,” or “how long does this task take on average,” you are no longer finding problems, you are estimating a metric, and small samples produce confidence intervals so wide they tell you nothing. Nielsen Norman Group’s own guidance is explicit: quantitative usability studies need 40 or more participants to produce ranges narrow enough to act on. Five users answering a quantitative question is not economical research. It is a guess wearing a lab coat.

A cleaner way to decide

Stop asking “how many users is enough” as if it has a fixed answer. Ask three questions in order.

First, am I looking for problems or measuring a number? Problems point you toward small qualitative rounds. Numbers point you toward samples of forty and up, or toward instrumented analytics instead of a test.

Second, how many distinct user groups touch this? Multiply your per-round count by the number of groups that genuinely behave differently. Do not average unlike users into one pool.

Third, what does a missed problem cost here? Low stakes and an early prototype justify five and a fast iteration. High stakes, a mature interface, or a one-way door on the build justify raising the floor toward ten or fifteen so a bad draw of participants cannot hide half your problems.

The five-user rule remains one of the most useful things ever written about research economics. It earned its fame. But it was an argument for testing early, testing often, and not overspending on a single round, and somewhere along the way it got flattened into a permission slip to test once and move on. The teams that get value out of it are the ones who kept reading past the title. If you want the practical companion to this, the mechanics of running a session that actually surfaces problems live in the prototype reality check, and the same sampling logic applies when you are deciding how many customer interviews it takes before you trust a pattern.

Five is the right answer to exactly one question. Your job is to make sure you are asking that question before you accept the answer.

Ty Sutherland

Ty Sutherland is the editor of Product Management Resources. With a quarter-century of product expertise under his belt, Ty is a seasoned veteran in the world of product management. A dedicated student of lean principles, he is driven by the ambition to transform organizations into Exponential Organizations (ExO) with a massive transformative purpose. Ty's passion isn't just limited to theory; he's an avid experimenter, always eager to try out a myriad of products and services. While he has a soft spot for tools that enhance the lives of product managers, his curiosity knows no bounds. If you're ever looking for him online, there's a good chance he's scouring his favorite site, Product Hunt, for the next big thing. Join Ty as he navigates the ever-evolving product landscape, sharing insights, reviews, and invaluable lessons from his vast experience.

Recent Posts