Experiments

We talked to 15 experimentation leaders about losing tests — here are their top takeaways

A graphic of a bar chart with an arrow pointing upward.

A losing test is not the opposite of a useful test. It is often the test that stops a confident team from scaling the wrong idea.

Experimentation programs love to celebrate wins. Wins make clean slides, support launch announcements, and translate easily into projected revenue. Losing and inconclusive tests are harder to narrate, even though they make up most of the work.

Across conversations with product, growth, data, and experimentation leaders, a different picture emerges. A negative result can protect revenue, expose a faulty assumption, reveal a segment that needs a different experience, or prevent months of engineering investment. The real failure is not a treatment losing. It is finishing the test without a trustworthy decision or reusable learning.

That view is consistent with the operating culture described in Harvard Business Review's account of Booking.com, where most ideas do not win and broad participation depends on making that outcome acceptable. It also follows the American Statistical Association's warning that a single threshold cannot substitute for scientific reasoning and full reporting.

The 15 leaders below work in different environments: marketplaces, financial services, ecommerce, AI products, enterprise software, logistics, and consumer products. Their shared lesson is simple: teams get more value from experimentation when they design for every possible result, not just the result they hope to announce.

What 15 leaders learned from tests that did not win

LeaderTeam or experienceLesson from losing tests
Medha UmarjiFanaticsA loss prevented can be as valuable as a winner shipped.
Dave MasseyUPSContext can turn the same friction from harmful to acceptable.
Kevin YangJPMorgan ChasePlanning for failure protects decisions before results create pressure.
Danielle OleanBoxA losing follow-up can reveal the boundary of a winning principle.
Dan LayfieldDiligent, Codecademy, Uber EatsAn inconclusive first attempt may contain the seed of a major win.
James FalzoneKargoA bad result is different from a badly designed experiment.
Pedro TabacofFinGuardrails reveal when an apparent AI improvement creates new harm.
Ilya IzrailevskyDoorDashMarketplace tradeoffs make one-sided wins dangerous.
Andrew WillinghamAtlassian, AmazonTesting assumptions is cheaper than defending expert intuition.
Fabian HansCogniteerFewer deep tests can teach more than a stream of shallow ones.
Kameron McNameeFyxerHigh velocity works when cheap losses feed the next iteration.
Nafis ShaikhChess.comSegment behavior can explain why an average result looks flat.
Ronny KohaviMicrosoft, Airbnb, AmazonSurprising lifts demand stronger scrutiny, not faster celebration.
Luke SonnetExperimentation leader and researcherFalse negatives matter when teams abandon valuable ideas too early.
Graham McNicollGrowthBookPrograms should count avoided harm and decisions, not wins alone.

These are not 15 versions of “failure is learning.” That phrase is too vague to guide a team. The useful patterns are more operational.

Treat avoided harm as measurable impact

Medha Umarji's Fanatics team runs close to 100 experiments a month. In her account of the program, only a minority become clear wins. Yet every losing treatment that would otherwise have shipped protects the baseline. Fanatics calls this the “do no harm” side of experimentation and uses non-inferiority guardrails when the goal is to improve an experience without materially damaging core business outcomes.

That framing changes the value calculation. If a proposed feature would have reduced conversion by 2%, the experiment did not create zero value merely because the treatment lost. It prevented the company from paying that 2% cost indefinitely. Fanatics is now working to make this avoided downside more visible alongside uplift from winners in its experimentation impact model.

Kevin Yang makes a similar case from JPMorgan Chase. The organization has attributed more than $1 billion in value to winning experiments, but he argues that losing tests may protect even more by preventing harmful rollouts. The practical move is to plan for failure before launch: decide what a negative effect means, who has authority to stop, and whether the treatment can be reversed. That avoids rewriting the decision rule after an executive-sponsored idea disappoints. The JPMorgan Chase interview makes avoided downside part of experimentation's business case.

Dave Massey's UPS story makes the same principle tangible. Requiring a recipient email in a shipping flow caused such a sharp conversion decline that the test stopped after about 24 hours. Years later, a team tested the field again for international shipments, this time explaining that customs might need to contact the recipient. The second version did not create the same problem. The input was not universally bad; unexplained friction was. That distinction came from testing the mechanism, not preserving the first loss as a blanket rule. UPS's broader program has linked experimentation to more than $500 million in incremental revenue, but its recipient-email lesson is about the losses the team was willing to stop.

Protect velocity from false wins

Use power planning, SRM checks, sequential testing, and multiple-comparison controls before a surprising result becomes a roadmap decision.

Read the Prevention Playbook

Separate a losing idea from a broken test

James Falzone asks Kargo teams to distinguish a “bad result” from a “bad experiment.” A treatment can lose even when assignment, exposure, metrics, duration, and analysis are sound. That is a useful result. A test with a sample-ratio mismatch, contaminated groups, broken tracking, or an implementation that does not represent the intended product change cannot support the same conclusion.

This distinction protects culture from empty positivity. Calling every result a win makes teams distrust the program. Kargo instead reviews failures in a biweekly retrospective and asks where the team failed, what context was missing, and what it would change. One model worked in direct demand but failed in third-party inventory because the environments differed. After the team learned that boundary and redesigned its approach, later iterations produced meaningful gains. The Kargo retrospective practice turns a negative result into a specific system improvement.

The technical checks matter because random noise will occasionally look extraordinary. Ronny Kohavi's work popularized Twyman's Law for online experiments: the more surprising the result, the more likely it is that something is wrong. Microsoft's experimentation research shows why teams need automatic assignment and data-quality checks at scale. Before interpreting a win or loss, inspect exposure, planned versus observed traffic, metric definitions, runtime, and confidence intervals. GrowthBook's A/A testing guidance and eBay's published work on automated sample-ratio-mismatch detection make these checks repeatable.

Graham McNicoll's practical standard is that a test should produce a decision, not merely a colored result card. If the data are invalid, the right outcome is “fix and rerun.” If the design was underpowered, it may be “collect more data” or “accept that this effect is below the decision threshold.” If the treatment is credibly harmful, it is “do not ship.” Those outcomes should not be collapsed into one failure bucket.

Investigate why averages hide the lesson

Danielle Olean's Box team learned that simplifying a pricing page helped, then learned that simplifying it further hurt. The loss did not invalidate the first finding. It revealed a boundary: customers benefited from reduced complexity until the page removed information they needed to choose confidently. Another cancellation-flow test won for an unexpected reason. Showing a cheaper plan made the customer's current plan feel more valuable—the team's “wine effect.” Box reports wins and losses in regular leadership rollups so the organization sees the actual learning rate, not a curated highlight reel. The Box ecommerce examples show why diagnostic metrics and qualitative interpretation matter.

Nafis Shaikh's work at Chess.com reinforces the need to inspect segments. A flat global result may combine a positive effect for new players with a negative effect for experienced ones, or hide differences across web and mobile. Segmentation is most useful when the groups are motivated before the result, have enough data, and lead to an actionable product choice. Fishing through dozens of segments after a loss simply creates another multiple-comparison problem. The Chess.com experimentation account is a reminder that user context belongs in the hypothesis.

Ilya Izrailevsky faces a more structural version at DoorDash. A marketplace change can help consumers while increasing Dasher wait time or merchant burden. A positive metric on one side is not automatically a business win. DoorDash's large program treats consumer, Dasher, and merchant outcomes as connected decision criteria. Its public engineering guidance likewise emphasizes design checks before analysis in a three-sided marketplace experimentation framework.

Give inconclusive ideas a disciplined second chance

Dan Layfield describes an early trial concept that looked inconclusive at Codecademy. The easy reaction would have been to close the document and move on. The team instead examined behavior, revised the proposition, and eventually found a version that produced roughly a 35% gain. The lesson is not to rerun every flat test until it wins. It is to preserve the mechanism-level evidence needed to decide whether the hypothesis deserves a materially different test. His Diligent interview argues against treating one implementation as the final verdict on an idea.

Luke Sonnet highlights the mirror image of false positives: false negatives. Low traffic, noisy metrics, implementation dilution, and inappropriate stopping rules can cause teams to discard a real improvement. Power analysis should begin with the smallest effect worth acting on. If the experiment cannot detect that effect in a useful period, the team needs a more sensitive metric, stronger treatment, variance reduction, or a different decision method. Spotify's Experiments with Learning framework similarly judges whether a test was capable of answering the intended question, not only whether it returned significance.

Fabian Hans argues for deeper experiments when learning is the goal. A rapid series of cosmetic changes can raise the launch count while leaving the underlying uncertainty untouched. A deep dive connects the intervention to a behavioral model, instruments the mechanism, and plans follow-ups. The Cogniteer conversation reframes velocity as time to useful understanding.

Make losses feed the next decision

Pedro Tabacof's team at Fin experiments on a non-deterministic AI support agent. A change can improve resolution while creating unacceptable promises or worse customer feedback. More conversation context once encouraged the agent to infer that it could issue refunds when it could not. The team revised the behavior and retested instead of accepting the apparent capability gain. Fin's production experimentation story shows why AI evals and A/B tests answer different questions: evals check known scenarios, while controlled production tests measure behavior across real traffic.

Kameron McNamee's small Fyxer team shows how losses can remain cheap. AI-assisted implementation and analysis help the team run hundreds of tests without requiring every treatment to become a major engineering project. High velocity works only if the experiment record makes the next iteration smarter. Otherwise the team is merely producing more discarded variants. Fyxer's 541-test year links speed to a repeatable idea-to-decision loop.

Andrew Willingham learned at Amazon and Atlassian that expertise does not remove uncertainty about user behavior. A product can make perfect sense to its builders and still confuse the people expected to use it. His framework converts assertions into assumptions, identifies the riskiest one, and chooses the cheapest credible test. The assumption-testing approach keeps a negative result from becoming a referendum on the person who proposed the idea.

Use a loss review that ends with an action

A practical review can fit on one page:

  1. Trust: Did assignment, exposure, metric computation, sample ratio, and runtime pass their checks?
  2. Decision: Did the treatment lose, remain inconclusive, or fail to meet a non-inferiority threshold?
  3. Mechanism: Which user behaviors changed, and do they explain the primary outcome?
  4. Heterogeneity: Were predeclared segments affected differently?
  5. Value protected: What likely downside did the test prevent from shipping?
  6. Next action: Ship the control, revise and rerun, target a segment, collect qualitative evidence, or stop investing.
  7. Memory: Where will the hypothesis, implementation, result, interpretation, and decision remain searchable?

Do not force every loss to generate another experiment. Sometimes the right learning is that the opportunity is too small, the mechanism is wrong, or the decision does not justify more traffic. The purpose of the review is to close the loop deliberately.

Mature teams also measure their program with more than win rate. Track trustworthy-decision rate, time to decision, losses caught before launch, repeated hypotheses avoided, and how often new briefs cite prior evidence. GrowthBook's experiment results support the statistical readout; the team still owns the causal interpretation and business decision.

The leaders in this group do not romanticize failure. They make it legible. They verify the test, quantify protected downside, examine the mechanism, and record what happens next. That is how a result nobody wanted becomes an asset the company can reuse.

Make every result trustworthy

See how experienced experimentation leaders combine guardrails, causal checks, and decision discipline before they call a test.

Read the Trustworthy Tests Recap

Table of Contents

Related Articles

See All Articles
Experiments

eCommerce experimentation: Insights and takeaways from the top companies

Aug 17, 2026
x
min read
Experiments

We talked to 4 leaders about getting a stuck experimentation team unstuck — here are their top takeaways

Aug 15, 2026
x
min read
Experiments

We talked to 10 leaders about building a culture of experimentation — here are their top takeaways

Aug 14, 2026
x
min read

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.