Most of Your Product Ideas Won’t Work. Plan For It.


Two-thirds of the ideas that shipped at Microsoft failed to move the metric they were built to improve. Ron Kohavi, who ran the experimentation platform there for years, has said it plainly: over two-thirds of ideas do not improve the metrics they were designed to improve. At Bing, the win rate dipped closer to 15%. At Airbnb, one analysis of 250 machine-learning ideas tested for search found that only 20 actually worked, a failure rate above 90%. Booking.com’s former design director Stuart Frisby put the number at 90% of tests fail, and added the part most people skip: “a lot of the low-hanging fruit has already been picked.”

Sit with those numbers, because they describe the companies that are best at this. These are teams with millions of daily users, dedicated experimentation platforms, statisticians on staff, and a decade of institutional practice. Their hit rate is roughly one in ten to one in three. Now ask yourself what the real success rate is at a team that ships features on gut feel, never instruments them, and moves on before anyone checks whether the metric moved.

The number nobody tells new product managers

When I was running IT operations, we treated deployments the way most product teams treat features: as things you build because you decided to build them. The decision was the event. Shipping was the finish line. Whether the thing actually did what we said it would do was a question that quietly disappeared under the next deadline. Nobody was lying. The feedback loop just never closed, so the belief that our ideas worked went permanently unchallenged.

Product management inherited the same blind spot, and it is more expensive here. A product manager’s entire job is to decide what gets built with scarce engineering time. If the honest base rate for “this idea will improve the metric” sits somewhere between 10% and 33% even at world-class shops, then a PM who greenlights ten confident bets a year should expect roughly two to three of them to land. The other seven or eight will be flat or negative.

That is not a sign of a bad product manager. It is the physics of the work. The problem is that almost nobody plans for it. Teams build roadmaps as if every item is a winner, staff them as if every winner will ship on time, and then act surprised when the quarter’s numbers barely move despite a full calendar of shipped work.

Why the win rate is so low (and why that is fine)

The low number is not a failure of imagination. It is what happens when you actually check.

Most product ideas sound good in a planning meeting because the room shares the same assumptions. The idea survives precisely because nobody in the room has the data to kill it. Kohavi’s famous example is the Bing ad-headline change from 2012: an engineer’s suggestion to lengthen the headline on paid results sat in the backlog for six months, rated low priority by experienced program managers. When someone finally ran the test, revenue jumped 12%, worth more than $100 million a year in the United States alone. The best revenue idea in Bing’s history was nearly killed by expert intuition, and only a test rescued it.

The mirror image happens far more often. Ideas everyone loves turn out to do nothing, or to quietly hurt a number no one was watching. That is the real value of a low win rate: it is evidence that the team is testing things it genuinely did not know the answer to. A team with a 90% win rate is not smarter. It is only testing things it already knew would work, which means it is learning almost nothing and probably leaving the big, non-obvious wins (the Bing headline) sitting in the backlog untouched.

The trap: treating a portfolio like a series of sure things

Here is where the math turns into a management problem. If your true hit rate is one in four, the worst thing you can do is pour six weeks of engineering into each unvalidated bet before you learn anything. Do that ten times and you have spent sixty engineer-weeks to find two or three winners, and you found them the slowest, most expensive way possible.

This is the mechanism behind what John Cutler named the feature factory back in 2016: teams praised for shipping volume, never for moving outcomes, with no feedback loop to tell them which of last quarter’s features actually mattered. A feature factory is not a team that ships bad ideas. It is a team that ships a normal mix of good and bad ideas and never finds out which was which, because it is always already building the next thing. The low win rate is invisible to them. They think they are batting .900 because they measure shipped, not worked.

The teams that internalize the real number do three things differently.

They shrink the bet before they place it. If most ideas fail, you want to fail on a fake door test or a rough prototype, not on a fully built, polished, six-week feature. The whole logic of a minimum viable product is a response to the base rate: spend the least possible to learn whether you are in the 25% or the 75%. Every dollar you spend before you have evidence is a dollar you are betting at unfavorable odds.

They instrument everything, before launch. You cannot have a win rate if you do not measure wins. The single most common reason a team does not know its real hit rate is that it never defined, in advance, what “worked” would mean for each feature. A number to move, a threshold, and a date to check it. Without that, every feature is a permanent maybe, and the roadmap becomes a list of things you did rather than things that worked.

They kill fast and say so out loud. Kohavi’s own summary of the discipline is blunt: fail fast, pivot fast. The organizations that experiment well have made killing an idea a normal, non-embarrassing event. That cultural piece matters more than the tooling. If shutting down your own feature is treated as a personal failure, nobody will do it, and the flat and negative bets will limp along consuming maintenance forever. This is the same gravity that makes teams keep shipping features they already suspect won’t work: the sunk cost is easier to defend than the empty column where the result should be.

What to do with this on Monday

You do not need a billion-variant testing platform like Booking’s to act on the base rate. Booking runs over a thousand concurrent tests and still watches 90% of them fail; the platform is not what makes them good, the honesty about the number is. You need three cheaper habits.

First, before you commit engineering time to anything sizeable, write down the metric it should move and the number that would count as a win. If you cannot name one, that is the finding: you are about to build on faith, and faith has a 25% hit rate.

Second, put a date on the wall for checking each shipped feature against its number. Not a retro, a specific “did this work” review. Most teams never schedule this, which is exactly why they never learn their real rate.

Third, keep a private tally of how many of your last ten shipped bets actually moved their metric. If the number is high, you are not winning; you are only testing the obvious and leaving the Bing headlines on the table. If the number looks like the pros, somewhere between two and four out of ten, you are doing the job correctly and the losses are the cost of finding the wins.

The product managers who last are not the ones with the highest hit rate. They are the ones who stopped being surprised by the low one, priced it into how they build, and made sure that when a winner did show up, they were spending small enough and measuring closely enough to actually catch it.


Sources: Ron Kohavi, “The Surprising Power of Online Experiments,” Harvard Business Review (2017); Ronny Kohavi interview, AB Tasty 1,000 Experiments Club; Trung Phan, “Booking: The $170B+ A/B Testing Machine”; John Cutler, “Scaled Feature Factories,” The Beautiful Mess.

Ty Sutherland

Ty Sutherland is the editor of Product Management Resources. With a quarter-century of product expertise under his belt, Ty is a seasoned veteran in the world of product management. A dedicated student of lean principles, he is driven by the ambition to transform organizations into Exponential Organizations (ExO) with a massive transformative purpose. Ty's passion isn't just limited to theory; he's an avid experimenter, always eager to try out a myriad of products and services. While he has a soft spot for tools that enhance the lives of product managers, his curiosity knows no bounds. If you're ever looking for him online, there's a good chance he's scouring his favorite site, Product Hunt, for the next big thing. Join Ty as he navigates the ever-evolving product landscape, sharing insights, reviews, and invaluable lessons from his vast experience.

Recent Posts