Experiments

We talked to 4 leaders about getting a stuck experimentation team unstuck — here are their top takeaways

A graphic of a bar chart with an arrow pointing upward.

A stalled experimentation program does not need a motivational campaign. It needs one constraint removed and one decision everyone can trust.

Teams get stuck in different places. One has a testing tool but no reliable exposure data. Another can launch treatments but waits weeks for an analyst. A third produces readouts that leadership ignores. A fourth runs many low-value tests because important questions cross organizational boundaries.

Advice from Medha Umarji at Fanatics, Dave Massey at UPS, Dan Layfield at Diligent, and James Falzone at Kargo points to a practical restart: diagnose the failure mode, choose a credible loop, make losses useful, and turn the first result into infrastructure for the next one.

Diagnose the kind of stuck

Review the last five ideas that should have become experiments. Mark the furthest stage each reached:

Failure modeEvidenceFirst intervention
No important questionsBacklog contains cosmetic variantsStart from customer and business uncertainty
No trustTeams dispute assignment or metric SQLRun A/A tests and publish validation
No implementation capacityApproved briefs wait for engineersAdd flag patterns, templates, and scoped self-service
No statistical pathTests cannot reach useful powerChange metric, treatment, population, or method
No decisionsReadouts sit without actionPredeclare owner and shipping criteria
No learning memoryTeams repeat old ideasCreate a searchable experiment record
No leadership permissionNegative executive ideas still shipEstablish evidence-based stop authority

Measure waiting time rather than relying on anecdotes. The constraint that appears most often deserves the first improvement cycle. GrowthBook's analysis of experiment velocity emphasizes that the target is important learning, not the largest launch count.

Medha Umarji: Change the executive behavior

Fanatics grew experimentation from a boutique conversion team to close to 100 tests a month. Medha Umarji credits executive engagement and humility: leaders inspect the data, question it, and accept when evidence contradicts their intuition.

That is the unblock when teams are afraid to test senior ideas or hide losses. A leader can restart the program by sponsoring a consequential reversible question, agreeing to the decision rule before launch, and reviewing the result publicly whether it wins or loses. One visible example does more than a general statement about being data-driven.

Fanatics also built an experiment wiki with causal interpretations, screenshots, and next steps that feed a future backlog. The Fanatics account shows how institutional memory turns a completed test into momentum.

Harvard Business Review's Booking.com case makes the same cultural connection: broad authority works because experiments are visible and people can challenge designs. Leadership creates permission and guardrails together.

Restart move: Ask an executive sponsor to precommit to acting on one upcoming result and share the full readout with the organization.

Rebuild trust in the test

Use power planning, SRM checks, sequential methods, and causal diagnostics to make the next decision defensible.

Read the Prevention Playbook

Dave Massey: Prove the loop on a business-critical surface

Dave Massey joined UPS when the company had a testing tool but not a durable program. His team received a pilot: demonstrate that customer-experience improvements could affect revenue. Removing distractions from the shipping flow produced an estimated $35 million annualized impact, but the data team still had to defend the result under scrutiny.

The pilot worked because the question mattered, the funnel was measurable, and the treatment was understandable. A trivial test might have launched faster but would not have earned the same organizational trust. The UPS story also shows the value of integrating UX research with behavioral measurement.

Microsoft researchers describe this dynamic as an experimentation adoption flywheel: a credible result increases belief, belief creates demand, demand supports infrastructure, and better infrastructure makes future results easier.

Restart move: Choose one high-traffic, reversible funnel problem with a committed owner and trusted outcome data. Publish the design and calculation, not only the uplift.

Dan Layfield: Reopen promising questions, not identical tests

Dan Layfield warns about moving on too quickly after an inconclusive result. At Codecademy, an early trial idea did not resolve, but the team inspected the behavior, changed the proposition, and later found a version with a much stronger result.

That does not justify rerunning every losing variant. The follow-up must contain a material change grounded in evidence. Ask whether the problem remains important, whether the treatment expressed the hypothesis strongly enough, and whether the test had the sensitivity to detect the minimum useful effect. The Diligent conversation turns an old result into a new design rather than a request for more luck.

Spotify's Experiments with Learning framework similarly asks whether a test began with clear intent and produced knowledge that affected a decision. That is a better restart metric than raw test count.

Restart move: Review three inconclusive or losing tests and select one whose mechanism still has evidence. Design a meaningfully different follow-up with a clear stop rule.

James Falzone: Make failure review operational

Kargo holds biweekly retrospectives that ask where the team failed. James Falzone separates a bad result from a bad experiment. A sound treatment can lose and still support a decision. A test with broken assignment or missing context cannot.

This direct language helps a stalled team because it avoids two unproductive extremes: blaming people for negative outcomes and celebrating every outcome as a win. The Kargo account describes how a model that failed in third-party inventory exposed missing context and informed later improvements.

Booking.com's paper on democratized experimentation connects this culture to a central repository, transparent data quality, and safeguards. Retrospectives work when the system preserves what changed afterward.

Restart move: Hold a 30-minute review with four fields: result integrity, mechanism learned, decision protected, and system change. Assign an owner to the system change.

A 30-day restart plan

Week 1: Establish the baseline

  • Map five blocked or completed tests across question, design, build, run, analysis, decision, and cleanup.
  • Name the repeated constraint.
  • Verify one exposure pipeline and one primary metric end to end.
  • Identify an executive sponsor and decision owner.

Week 2: Design one credible test

  • Select an important, reversible question with sufficient traffic.
  • Predeclare the hypothesis, assignment, metric, guardrails, duration, and outcomes.
  • Run an A/A or shadow validation if the surface is new.
  • Create the implementation behind a feature flag with a rollback path.

DoorDash's engineering guidance shows how design validation reduces later toil in its experimentation framework.

Week 3: Launch visibly

  • Publish the brief where peers can review it.
  • Ramp from a small safe audience while checking exposure and guardrails.
  • Monitor data quality without interpreting immature outcome movement.
  • Record questions and corrections in the experiment timeline.

Week 4: Close the loop

  • Present the result, interval, diagnostics, limitations, and decision.
  • State realized or prevented business impact conservatively.
  • Roll out or revert; assign code and flag cleanup.
  • Add the record to a searchable repository.
  • Choose the next workflow improvement based on observed waiting time.

GrowthBook's experimentation platform supports reusable metrics, analysis, feature delivery, and program visibility. The operating discipline still belongs to the team.

Do not restart with a volume quota

A demand for 20 tests next quarter can create motion while preserving the original dysfunction. Teams split ideas into smaller variants, select high-traffic surfaces regardless of strategic value, and avoid difficult cross-functional questions.

Track time to trustworthy decision, eligible changes tested, result integrity, prior evidence reused, avoided losses, and closure time. eBay's work on automated randomization validation is a reminder that scale depends on quality checks becoming part of the platform, not on relaxing them.

A stuck program becomes unstuck when the organization believes one complete loop. Find the constraint, resolve an important uncertainty, honor the result, preserve the learning, and make the second loop cheaper than the first.

Design for long-term learning

Hear how experienced leaders approach metric choice, trustworthy setup, failure planning, and organizational adoption.

Watch the Leadership Session

Table of Contents

Related Articles

See All Articles
Experiments

eCommerce experimentation: Insights and takeaways from the top companies

Aug 17, 2026
x
min read
Experiments

We talked to 15 experimentation leaders about losing tests — here are their top takeaways

Aug 14, 2026
x
min read
Experiments

We talked to 10 leaders about building a culture of experimentation — here are their top takeaways

Aug 14, 2026
x
min read

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.