We talked to 6 growth leaders about turning experimentation into a revenue engine — here are their top takeaways

A revenue engine is not a dashboard of winning tests. It is a loop that repeatedly finds, measures, ships, and compounds product improvements.
One large win can establish credibility, but it does not create a program. Revenue becomes repeatable when the organization can move from an important question to a trustworthy decision, roll the treatment out, verify the effect, and turn the learning into the next hypothesis.
Six growth leaders—from UPS, Fanatics, JPMorgan Chase, Box, DoorDash, and Fyxer—show the components of that loop.
The six components
| Leader | Organization | Engine component |
|---|---|---|
| Dave Massey | UPS | Start with a measurable business surface and defend the result. |
| Medha Umarji | Fanatics | Convert completed tests into a future backlog. |
| Kevin Yang | JPMorgan Chase | Count prevented losses alongside winning impact. |
| Danielle Olean | Box | Explain the customer mechanism behind revenue movement. |
| Ilya Izrailevsky | DoorDash | Balance revenue across a connected marketplace. |
| Kameron McNamee | Fyxer | Lower the marginal cost of every learning cycle. |
1. UPS: Establish a credible unit of impact
Dave Massey's early UPS pilot applied ecommerce thinking to the shipping flow. Removing distracting navigation increased conversion with an estimated $35 million annualized impact. The data team then defended the calculation through detailed review.
That review established a unit the company could trust: eligible shipping sessions, a randomized treatment, a measured conversion difference, and a projection tied to the actual funnel. Years of similar work have contributed to more than $500 million in attributed incremental revenue according to the UPS experimentation account.
Microsoft's research on controlled experimentation at scale makes the wider point: reliable causal evidence improves resource allocation as well as individual features.
Engine rule: Store the unit, eligible population, absolute effect, uncertainty interval, rollout, and projection window beside every revenue claim.
2. Fanatics: Make each result create the next opportunity
Medha Umarji's Fanatics team grew from roughly 10 tests a month to close to 100. Its experiment wiki stores results, causal interpretations, visual evidence, and next steps. Recommendations feed a future backlog, often using capabilities already built for an earlier test.
This is how experimentation compounds. The first treatment pays for learning and implementation; the follow-up reuses both. Meta-analysis across three or more related tests helps the team see patterns that a single result cannot establish. The Fanatics program attributes a meaningful share of annual growth to this system while still rerunning suspicious wins.
Booking.com's paper on democratized experimentation similarly identifies a central repository of successes and failures as infrastructure for organization-wide learning.
Engine rule: No experiment closes without a causal interpretation, next action, and searchable record.
Turn results into business impact
Build primary, revenue, margin, diagnostic, and guardrail metrics that make experiment decisions auditable.
Read the KPI Playbook3. JPMorgan Chase: Add the value of not shipping harm
Kevin Yang's organization has estimated more than $1 billion in value from winning experiments and supports roughly 100 product teams. He argues that negative tests may protect even more value by preventing damaging rollouts.
The JPMorgan Chase interview recommends planning for failure before launch. That removes pressure to reinterpret a negative result when a senior stakeholder wants the feature.
Prevented loss should remain a separate reporting category. Use the observed treatment downside, credible eligible rollout, and conservative duration. Do not treat it as recognized revenue. The discipline makes avoided harm visible without turning every stopped idea into an inflated savings claim.
Engine rule: Report realized uplift, projected uplift, prevented downside, and cost savings separately.
4. Box: Require a plausible revenue mechanism
Danielle Olean's Box ecommerce team follows top-line movement through diagnostic behavior. A cheaper-plan offer in cancellation worked partly by making the existing plan appear more valuable. One pricing-page simplification helped; additional simplification hurt.
The Box examples demonstrate why an experiment engine needs explanatory metrics. If revenue moves but product views, cart behavior, plan selection, or retention do not move as expected, the team should investigate before generalizing the rule.
Research on ecommerce metrics in experiments also warns that transaction-based ratio metrics require careful treatment because purchases and items are not independent observations.
Engine rule: Every revenue result includes the expected behavioral chain and evidence for or against it.
5. DoorDash: Optimize the system, not one side
Ilya Izrailevsky's DoorDash program operates across consumers, Dashers, and merchants. A treatment that raises order revenue can still reduce contribution profit, increase wait time, or degrade merchant operations.
DoorDash uses designs that fit marketplace interference, including switchback experimentation. Its business-policy research reports a fractional-factorial approach for evaluating interacting promotion components more efficiently.
The lesson extends beyond marketplaces. Revenue is a gross outcome; the engine needs contribution margin, operational cost, customer experience, and retention guardrails wherever the treatment can change them.
Engine rule: Match randomization and the scorecard to the economic system that produces revenue.
6. Fyxer: Reduce cost per decision
Kameron McNamee's four-person Fyxer growth team ran 541 experiments in a year using AI-assisted coding, analysis, and GrowthBook. The point is not that every team needs that count. It is that lower implementation and analysis cost expands the number of economically sensible questions.
The Fyxer account shows how small teams can connect product-led growth and experimentation when the workflow is reusable. AI drafts treatments and analyses; humans still own the hypothesis, verification, and decision.
DORA's research on software delivery performance provides a useful parallel: small, frequent, recoverable changes improve the delivery system. In experimentation, flags, automated checks, and governed metrics make learning increments smaller and safer.
Engine rule: Measure total time and specialist effort per trustworthy decision, then automate the repeated steps.
Operate the engine as a portfolio
Balance easy, medium, and hard tests. Easy tests maintain cadence and improve known funnels. Medium tests probe stronger behavior changes. Hard tests address pricing, ranking, onboarding, marketplace policy, and new value propositions.
Use a portfolio scorecard:
- Eligible product changes tested.
- Time from question to trustworthy decision.
- Decision rate and invalid-test rate.
- Realized and projected incremental value.
- Prevented downside.
- Strategic surfaces covered.
- Prior experiments cited in new briefs.
- Cleanup time after decision.
Validate the sum with a program-level holdout when the opportunity cost is justified. GrowthBook's holdout framework estimates the cumulative effect of shipped changes and the actual exposure path rather than adding independent winner estimates.
Build the minimum viable engine
Start with one cross-functional team and one measurable surface. Create stable assignment and exposure, five governed metrics, an experiment brief, automatic data-quality checks, a predeclared decision owner, and a closure checklist. GrowthBook's experiment platform can provide the shared analysis and feature flag delivery needed to reuse the loop.
After three completed tests, hold a meta-review. Which stage waited? Which metric failed? Which idea should be iterated? Which claim can finance reproduce? Improve that system before expanding participation.
Revenue engines are built from disciplined repetition. The six leaders make the same tradeoff visible: move quickly enough to ask more valuable questions, but keep the evidence strong enough that product and finance will act on the answer.
Keep the portfolio trustworthy
Review the causal and statistical checks leaders use before individual wins become a program-level impact claim.
Read the Trustworthy Tests RecapRelated Articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics—free.


