Experiments

We talked to 6 growth leaders about turning experimentation into a revenue engine — here are their top takeaways

A graphic of a bar chart with an arrow pointing upward.

A revenue engine is not a dashboard of winning tests. It is a loop that repeatedly finds, measures, ships, and compounds product improvements.

One large win can establish credibility, but it does not create a program. Revenue becomes repeatable when the organization can move from an important question to a trustworthy decision, roll the treatment out, verify the effect, and turn the learning into the next hypothesis.

Six growth leaders—from UPS, Fanatics, JPMorgan Chase, Box, DoorDash, and Fyxer—show the components of that loop.

The six components

LeaderOrganizationEngine component
Dave MasseyUPSStart with a measurable business surface and defend the result.
Medha UmarjiFanaticsConvert completed tests into a future backlog.
Kevin YangJPMorgan ChaseCount prevented losses alongside winning impact.
Danielle OleanBoxExplain the customer mechanism behind revenue movement.
Ilya IzrailevskyDoorDashBalance revenue across a connected marketplace.
Kameron McNameeFyxerLower the marginal cost of every learning cycle.

1. UPS: Establish a credible unit of impact

Dave Massey's early UPS pilot applied ecommerce thinking to the shipping flow. Removing distracting navigation increased conversion with an estimated $35 million annualized impact. The data team then defended the calculation through detailed review.

That review established a unit the company could trust: eligible shipping sessions, a randomized treatment, a measured conversion difference, and a projection tied to the actual funnel. Years of similar work have contributed to more than $500 million in attributed incremental revenue according to the UPS experimentation account.

Microsoft's research on controlled experimentation at scale makes the wider point: reliable causal evidence improves resource allocation as well as individual features.

Engine rule: Store the unit, eligible population, absolute effect, uncertainty interval, rollout, and projection window beside every revenue claim.

2. Fanatics: Make each result create the next opportunity

Medha Umarji's Fanatics team grew from roughly 10 tests a month to close to 100. Its experiment wiki stores results, causal interpretations, visual evidence, and next steps. Recommendations feed a future backlog, often using capabilities already built for an earlier test.

This is how experimentation compounds. The first treatment pays for learning and implementation; the follow-up reuses both. Meta-analysis across three or more related tests helps the team see patterns that a single result cannot establish. The Fanatics program attributes a meaningful share of annual growth to this system while still rerunning suspicious wins.

Booking.com's paper on democratized experimentation similarly identifies a central repository of successes and failures as infrastructure for organization-wide learning.

Engine rule: No experiment closes without a causal interpretation, next action, and searchable record.

Turn results into business impact

Build primary, revenue, margin, diagnostic, and guardrail metrics that make experiment decisions auditable.

Read the KPI Playbook

3. JPMorgan Chase: Add the value of not shipping harm

Kevin Yang's organization has estimated more than $1 billion in value from winning experiments and supports roughly 100 product teams. He argues that negative tests may protect even more value by preventing damaging rollouts.

The JPMorgan Chase interview recommends planning for failure before launch. That removes pressure to reinterpret a negative result when a senior stakeholder wants the feature.

Prevented loss should remain a separate reporting category. Use the observed treatment downside, credible eligible rollout, and conservative duration. Do not treat it as recognized revenue. The discipline makes avoided harm visible without turning every stopped idea into an inflated savings claim.

Engine rule: Report realized uplift, projected uplift, prevented downside, and cost savings separately.

4. Box: Require a plausible revenue mechanism

Danielle Olean's Box ecommerce team follows top-line movement through diagnostic behavior. A cheaper-plan offer in cancellation worked partly by making the existing plan appear more valuable. One pricing-page simplification helped; additional simplification hurt.

The Box examples demonstrate why an experiment engine needs explanatory metrics. If revenue moves but product views, cart behavior, plan selection, or retention do not move as expected, the team should investigate before generalizing the rule.

Research on ecommerce metrics in experiments also warns that transaction-based ratio metrics require careful treatment because purchases and items are not independent observations.

Engine rule: Every revenue result includes the expected behavioral chain and evidence for or against it.

5. DoorDash: Optimize the system, not one side

Ilya Izrailevsky's DoorDash program operates across consumers, Dashers, and merchants. A treatment that raises order revenue can still reduce contribution profit, increase wait time, or degrade merchant operations.

DoorDash uses designs that fit marketplace interference, including switchback experimentation. Its business-policy research reports a fractional-factorial approach for evaluating interacting promotion components more efficiently.

The lesson extends beyond marketplaces. Revenue is a gross outcome; the engine needs contribution margin, operational cost, customer experience, and retention guardrails wherever the treatment can change them.

Engine rule: Match randomization and the scorecard to the economic system that produces revenue.

6. Fyxer: Reduce cost per decision

Kameron McNamee's four-person Fyxer growth team ran 541 experiments in a year using AI-assisted coding, analysis, and GrowthBook. The point is not that every team needs that count. It is that lower implementation and analysis cost expands the number of economically sensible questions.

The Fyxer account shows how small teams can connect product-led growth and experimentation when the workflow is reusable. AI drafts treatments and analyses; humans still own the hypothesis, verification, and decision.

DORA's research on software delivery performance provides a useful parallel: small, frequent, recoverable changes improve the delivery system. In experimentation, flags, automated checks, and governed metrics make learning increments smaller and safer.

Engine rule: Measure total time and specialist effort per trustworthy decision, then automate the repeated steps.

Operate the engine as a portfolio

Balance easy, medium, and hard tests. Easy tests maintain cadence and improve known funnels. Medium tests probe stronger behavior changes. Hard tests address pricing, ranking, onboarding, marketplace policy, and new value propositions.

Use a portfolio scorecard:

  • Eligible product changes tested.
  • Time from question to trustworthy decision.
  • Decision rate and invalid-test rate.
  • Realized and projected incremental value.
  • Prevented downside.
  • Strategic surfaces covered.
  • Prior experiments cited in new briefs.
  • Cleanup time after decision.

Validate the sum with a program-level holdout when the opportunity cost is justified. GrowthBook's holdout framework estimates the cumulative effect of shipped changes and the actual exposure path rather than adding independent winner estimates.

Build the minimum viable engine

Start with one cross-functional team and one measurable surface. Create stable assignment and exposure, five governed metrics, an experiment brief, automatic data-quality checks, a predeclared decision owner, and a closure checklist. GrowthBook's experiment platform can provide the shared analysis and feature flag delivery needed to reuse the loop.

After three completed tests, hold a meta-review. Which stage waited? Which metric failed? Which idea should be iterated? Which claim can finance reproduce? Improve that system before expanding participation.

Revenue engines are built from disciplined repetition. The six leaders make the same tradeoff visible: move quickly enough to ask more valuable questions, but keep the evidence strong enough that product and finance will act on the answer.

Keep the portfolio trustworthy

Review the causal and statistical checks leaders use before individual wins become a program-level impact claim.

Read the Trustworthy Tests Recap

Table of Contents

Related Articles

See All Articles
Experiments

eCommerce experimentation: Insights and takeaways from the top companies

Aug 17, 2026
x
min read
Experiments

We talked to 4 leaders about getting a stuck experimentation team unstuck — here are their top takeaways

Aug 15, 2026
x
min read
Experiments

We talked to 15 experimentation leaders about losing tests — here are their top takeaways

Aug 14, 2026
x
min read

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.