Fit Is a Stack of Hypotheses
Dan Olsen splits product-market fit into five testable layers under a six-step process. Test clickable mockups on waves of five users before writing code, then read the verdict in cohort retention.
The Core Insight
Intuit launched Quicken into a market that already held 46 personal finance products. Research convinced the founders that none of the 46 had fit. They bet on a checkbook design that matched how buyers thought, tested every version on users, and rode ease of use to number one. The founders joked afterward about 47th mover advantage.
The Lean Product Playbook is Dan Olsen's 2015 manual for making that outcome a process. His diagnosis is flat: products die when they do not meet customer needs better than the alternatives. His response treats fit as an engineering problem, decomposed into hypotheses and tested in the cheapest medium that can kill each one.
Most product teams treat fit as a verdict the market returns after launch, so they build first and learn from shipping. Olsen argues the verdict decomposes in advance. Five hypotheses stack under every product, and a clickable mockup tests most of them. A team working bottom-up finds the broken layer before paying for code.
He earned the method in production. He studied Lean manufacturing at Virginia Tech, led Quicken to record sales, ran product at Friendster, and consulted for Facebook and Box.
The Framework
The Product-Market Fit Pyramid stacks five hypothesis layers. Target customer and underserved needs sit at the base and define the market. Value proposition, feature set, and user experience sit above and define the product. Fit completes the pyramid as the seam between the two groups, measuring how well the top three layers satisfy the bottom two.
The Lean Product Process walks the stack in six steps.
- Determine your target customers.
- Identify underserved customer needs.
- Define your value proposition.
- Specify your MVP feature set.
- Create your MVP prototype.
- Test your MVP with customers.
Settle the first three steps before looping on the last three. Fix errors at the lowest layer first, because a low error invalidates every layer above it. Iterating on UX cannot save a wrong target customer.
Needs live in problem space, and products live in solution space. Customers give vague feedback on abstract benefits and sharp feedback on concrete artifacts, so the fastest problem-space learning comes from showing solution-space artifacts. At Intuit, Cook named TurboTax's biggest competitor as pen and paper, since more Americans prepared taxes by hand than used all tax software combined.
Key Ideas
The Upper Left Quadrant Is the Opportunity
Olsen plots every need on two axes, importance vertical and satisfaction with current alternatives horizontal. Importance gets a 5-point unipolar scale, satisfaction a 7-point bipolar one. Low importance is dead on either side. High importance with high satisfaction is a served market you must displace. High importance with low satisfaction is underserved, and underserved needs are the business.
Uber lived in that quadrant. Taxis mattered and disappointed, and Uber answered each complaint with a mechanism: live car maps, driver ratings, fare estimates, stored cards. Leaked December 2013 data showed over 400,000 active clients, over 800,000 rides per week, and a run rate above 1 billion dollars a year. A year later Uber raised 1.2 billion dollars at a 40 billion dollar valuation.
Importance minus satisfaction gives the gap, and the gap prices a 5-point shortfall the same at every importance, which misleads. Ulwick's opportunity score adds importance back: importance plus the gap, floored at zero, on 0 to 10 scales. Above 15 is very attractive, below 10 is unattractive. Importance 10 with satisfaction 5 scores 15, while importance 6 with satisfaction 1 scores 11.
Olsen's own version multiplies. Value delivered is importance times satisfaction, and opportunity is importance times the satisfaction shortfall. A need at 70 percent importance and 70 percent satisfaction leaves 0.21 of opportunity. A need at 90 percent importance and 30 percent satisfaction leaves 0.63, three times as much. In one 13-feature survey, feature X sat at 82 percent importance and 55 percent satisfaction, an opportunity of 0.37, with 11 features under 0.25.
Win One Performance Benefit and One Delighter
Kano sorts value proposition rows into three kinds. Must-haves create no satisfaction when met and kill you when missed. Performance benefits scale, where more is better. Delighters surprise, and nobody misses them when absent. Needs migrate: GPS navigation entered cars as a mid-1990s delighter and hardened into an expectation.
The value proposition table scores those rows for you and every alternative, indirect ones included. The winning shape: match all must-haves, win one chosen performance benefit, and carry a unique delighter. Search engines competed on result count, freshness, and relevance. Count and freshness commoditized, Google won relevance with PageRank, and its Instant delighter saved a measured two to five seconds per search.
Cuil ran the counter-case. It launched in 2008 against a Google holding over 60 percent share, claiming an index of 120 billion pages, an estimated three times Google's. Response time and relevance came in poor, the differentiators bought nothing on top of a losing core, and the engine closed after two years. The rule: match the incumbent on the important performance benefits before differentiating.
Scope the MVP by Return on Effort
An MVP carries the minimum functionality the target customer considers viable, and viable is the customer's call. Olsen scopes it with division. Break features into chunks small enough to estimate, score each chunk's customer value on a ratio scale, estimate developer-weeks, and divide. In his grid an idea worth 6 at 2 weeks returns 3, and an idea worth 6 at 4 weeks returns 1.5. Ties go to the smaller chunk, because smaller ships sooner and teaches sooner.
Assembling the MVP candidate takes three rules. Include every must-have. Add enough chunks of the main performance benefit that customers see the difference. Add the top delighter unless the performance lead is already wide. The leftmost column becomes v1 and the rest sketch a roadmap. He warns against planning past one or two minor versions, because everything changes once customers see the prototype.
Test in Waves of Five, One Hour Each
Olsen replaces the MVP argument with MVP tests, sorted on a two by two: product against marketing, qualitative against quantitative. Start qualitative, because teams that jump to counting cannot explain their counts. Artifacts vary on fidelity and interactivity, running from grayscale wireframes through clickable mockups to interactive prototypes and live product. Manual hacks count too. Airbnb tested professional listing photography by recruiting photographers and uploading shots by hand, and photographed listings booked two to three times the market average.
Test one customer at a time, in waves of five to eight, enough to surface the major issues. Sessions run about an hour, give or take 15 minutes. About 10 percent of recruits no-show, so schedule one extra. The rule that saves the practice is blind scheduling: book a recurring slot, three users every Tuesday afternoon, before knowing what you will show.
Moderation has one law: leading questions corrupt results faster than anything else. Ask open questions, echo actions back before probing, and tolerate silence. Never rescue a struggling user, because nobody holds hands after launch. Close each session with 0 to 10 ratings on value and ease. Progress across waves shows three ways: more positives, rising ratings, and silence where old complaints lived.
A clean run is weak evidence of value. On Olsen's news product, wave three users sailed through and around 20 percent still refused to use it. The blocker sat at the pyramid's base: three distinct ways people prefer to get news, and a product built for one. The team redefined the target customer instead of the interface.
Iterate by Layer, Pivot by Mountain
Olsen amends build-measure-learn into hypothesize-design-test-learn, since clickable wireframes need no build and observation needs no dashboard. Each wave's issues map onto pyramid layers, and repairs run bottom-up.
His worked table converges over four waves of five users. Complaints about a missing feature fell from 80 percent to zero once the feature shipped. Registration difficulty took two redesigns, falling from 60 to 40 to zero percent. Median value ratings climbed from 7 to 9 and ease from 5 to 9. New complaints surfaced mid-loop as each fix exposed the next seam, which is the loop working.
A pivot changes a bottom hypothesis, the target customer or the differentiators, and a UX tweak does not qualify. The triggers: several waves without gains, lukewarm targets, and no customer archetype excited about the MVP. Flickr began as an online game, kept the photo tool, launched in February 2004, and sold to Yahoo in March 2005. Instagram cut Burbn down to photo, comment, and like for its October 2010 launch, and Facebook paid about 1 billion dollars in April 2012.
The end-to-end case ran the loop for a client. Two marketing-data concepts went in front of eight customers each, in 90-minute evening sessions with personalized printed mockups. Neither concept compelled anyone, and the invented marketing score confused users. The team pivoted to junk-mail blocking, rebuilt the mockups around 31 mail types in seven categories, and retested. Every second-round customer asked to be notified at launch, a request no first-round customer made. The whole project took under two months.
Retention Is the Scoreboard
After launch, surveys carry two standards. Net promoter score subtracts detractors, scores 0 through 6, from promoters, scores 9 and 10, on the recommendation question. The Sean Ellis question asks how users feel if the product disappears, and 40 percent or more answering very disappointed marks probable fit. Send it to a random sample of recent repeat users.
Retention outranks both, because it is the one metric not conflated with acquisition or conversion. Group users into cohorts, plot the percentage active against days since signup, and read three parameters: initial drop, descent rate, terminal value. Terminal value decides. A product flattening at 50 percent beats one flattening at 1 percent before you know anything else about either. A/B testing comes afterward, because bottom layers harden like tectonic plates once built.
The equation of the business peels revenue into levers. Profit is revenue minus cost. Lifetime value is revenue per user times average lifetime times gross margin. Average lifetime is one over churn, so 5 percent monthly churn means a 20-month customer. Successful SaaS holds lifetime value above three times acquisition cost. For a new product, work retention first, then conversion, then acquisition, because prospects pushed through a leaky funnel are wasted.
Friendster shows the machine at speed. Olsen decomposed viral growth into five ratio metrics whose product is the viral coefficient, and measured baselines. Fifteen percent of users sent invites, senders averaged 2.3 invites each, and registration converted at 85 percent. Headroom math picked the target: registration held 18 percent of upside, invites per sender held 6,520 percent, taking Dunbar's 150 relationships as the ceiling. An address book importer for Yahoo Mail took one week from one product manager and one developer. The metric settled around 5.3, a 2.3 times gain, worth 2.3 times the new customers from viral growth.
Practical Applications
Write the five hypotheses on one page: customer, underserved needs, value proposition, v1 features, experience. State each so a single test can kill it.
Rate your benefit list with target customers on the importance and satisfaction scales. Multiply importance by the satisfaction shortfall and build where the number is largest.
Book the recurring slot before any artifact exists: five customers, one afternoon, every week, screened by a short survey. Show clickable mockups for about an hour each, ask open questions, and rescue nobody. Revise between waves, never during one.
Map each complaint to its pyramid layer and fix the lowest first. Treat several flat waves with no excited archetype as the pivot trigger, and change the hypothesis rather than the wireframe.
Once live, chart cohort retention on days since signup and read the terminal value monthly. Send the Ellis question to a random sample of repeat users and hold the answers against the 40 percent line.
Who This Is For
Founders between idea and first release get the most, because every chapter assumes change is still cheap. Product managers on live products get the analytics half: retention curves, metric order, the equation of the business. Skip it if your curve already flattens high and your question is growth, because the book stops where scaling starts.
The book is a practitioner manual, and it reads like one. The worked tables are stylized illustrations rather than audited data, and the Airbnb photography result appears in two conflicting versions in the text. The tools overlap competing frameworks: the loop amends Ries, the opportunity score is Ulwick's, the pirate metrics are McClure's, and Kano arrives whole. The 2015 tool lists were dated on arrival, and the Agile chapter is a generic primer.
The Decision
Run the wave test this week. Write a short screener for your target customer, and recruit five of them for one afternoon. Book the slot before the mockup is finished. Hold each session for an hour: open questions, no rescues, a 0 to 10 value rating at the end.
Five sessions will name the broken layer. The calendar invite goes out today.