We talked to 10 leaders about building a culture of experimentation — here are their top takeaways

Experimentation becomes culture when the organization changes how it handles uncertainty—not when it buys a testing tool.
The phrase “culture of experimentation” is easy to flatten into encouragement to run more A/B tests. Leaders who have built real programs describe something more demanding. Teams need permission to challenge intuition, a trustworthy way to test, shared language for evidence, and a habit of recording what the company learned.
We reviewed conversations with 10 leaders who have built or operated experimentation programs at Fanatics, UPS, DoorDash, JPMorgan Chase, Box, Diligent, Kargo, Fyxer, Cogniteer, Amazon, and Atlassian. Their companies differ in traffic, regulation, product, and maturity. Yet their advice converges on 6 operating choices.
The 10 leaders and the cultural behavior each emphasizes
| Leader | Organization or experience | Cultural behavior |
|---|---|---|
| Medha Umarji | Fanatics | Executives model humility when data contradicts intuition. |
| Dave Massey | UPS | Earn trust by connecting rigor to customer and revenue outcomes. |
| Ilya Izrailevsky | DoorDash | Make results visible across product teams and marketplace sides. |
| Kevin Yang | JPMorgan Chase | Plan for losing outcomes before pressure changes the decision. |
| Danielle Olean | Box | Report wins and losses to leadership on a regular cadence. |
| Dan Layfield | Diligent, Codecademy, Uber Eats | Preserve promising mechanisms after inconclusive first attempts. |
| James Falzone | Kargo | Review failures directly without calling every outcome a win. |
| Kameron McNamee | Fyxer | Remove implementation friction so small teams can learn quickly. |
| Fabian Hans | Cogniteer | Reward depth of learning, not only number of launches. |
| Andrew Willingham | Amazon, Atlassian | Convert opinions into testable assumptions instead of identity battles. |
Leadership must make it safe to be wrong
Medha Umarji credits Fanatics' experimentation culture partly to executives who want to see the data and are willing to change their minds. That behavior matters more than a poster about curiosity. When a CEO publicly accepts that a favored idea lost, everyone else learns that evidence can outrank seniority. The conversation shifts from “why test this?” to “how should we test this?”
Fanatics' program grew from roughly 10 monthly tests to close to 100, and experimentation contributes materially to annual growth. But Medha's cultural point is not the number. It is the humility that lets a result alter a roadmap. The Fanatics program account also describes an experiment wiki that turns completed work into future hypotheses.
Andrew Willingham uses a related technique: state an assumption instead of defending a conclusion. “Users will understand this workflow” is easier to test than “this is the right design.” The wording separates a person's status from the result. His Amazon-to-Atlassian framework helps teams locate the riskiest belief before investing in a full implementation.
This is consistent with the well-documented Booking.com model. Harvard Business Review describes broad launch authority paired with transparent experiment proposals and the ability for colleagues to challenge unsafe work. Psychological permission is paired with operational visibility.
Trust grows when rigor is visible
Dave Massey's first major UPS experiment removed distractions from a shipping flow and produced an estimated $35 million annualized conversion impact. The number did not create instant acceptance. The data team had to defend assignment, measurement, and interpretation under intense scrutiny. That process helped establish the program's credibility.
UPS later connected UX research and experimentation under one organization. Behavioral data shows what changed; customer research helps explain why. When a treatment fails, the team can return to customers rather than guessing at a mechanism. The UPS experimentation story shows how rigor and customer understanding reinforce each other.
Trustworthy culture therefore requires more than statistical training. It needs stable assignment, exposure data, reviewed metric definitions, sample-ratio-mismatch checks, and decision rules. Booking.com's technical paper on democratizing controlled experiments explicitly connects culture to safeguards, transparent data quality, a central repository, and loose coupling between experiments and business logic.
Build trust into every test
Review the practical checks that protect a growing experimentation program from false wins, broken assignment, and premature decisions.
Read the Prevention PlaybookMake learning visible enough to compound
Danielle Olean's Box team sends recurring impact reports and leadership rollups that include negative and positive outcomes. That prevents survivorship bias. If executives only hear about winners, they develop an unrealistic view of win rate and may punish the next team whose reasonable idea does not move the metric. The Box ecommerce conversation illustrates why causal interpretations belong beside topline results.
Ilya Izrailevsky describes broad readouts at DoorDash, where experimentation spans consumers, Dashers, and merchants. Visibility helps a result from one product surface inform another and makes marketplace tradeoffs harder to ignore. DoorDash's engineering team treats velocity, analyst toil, rigor, and compute cost as a combined problem in its experimentation framework.
An experiment repository should preserve more than a screenshot and uplift estimate. Record the decision context, hypothesis, assignment unit, exposure rule, primary metric, guardrails, segments, implementation, result, interpretation, and next action. Searchability matters because organizational learning decays when only the original analyst remembers the caveat.
Microsoft researchers describe adoption as a flywheel: early trustworthy examples increase belief, belief creates demand, demand justifies better infrastructure, and easier workflows create more examples. Their multi-company experimentation flywheel is a better cultural model than a one-time training campaign.
Normalize losses without lowering the bar
James Falzone's Kargo retrospectives ask teams where they failed. He distinguishes a sound experiment with an unfavorable result from a broken experiment that cannot support a conclusion. That keeps “failure is learning” from becoming an excuse for weak design. The Kargo story shows that a context-specific loss can guide a redesign and later gains.
Kevin Yang recommends planning for failure before launch. If a treatment hurts the primary metric, violates a guardrail, or produces an inconclusive result, the team should already know the default action and escalation path. This is particularly important in financial services, where pressure from senior sponsors can collide with risk obligations. His JPMorgan Chase interview treats avoided harm as a central source of value.
Dan Layfield warns against a different mistake: interpreting the first inconclusive implementation as proof that the underlying opportunity is worthless. At Codecademy, a trial concept that initially failed to resolve later informed a substantially stronger version. The discipline is to revise the mechanism, not rerun the same weak treatment until random variation produces a win. His Diligent conversation makes follow-up quality part of culture.
Remove friction without removing responsibility
Kameron McNamee's four-person growth team at Fyxer used AI-assisted coding and analysis to run 541 tests in a year. That volume was possible because the path from idea to implementation was short. The team did not need a new cross-functional project for every copy, onboarding, or product change. The Fyxer case demonstrates how lowering marginal cost expands the set of questions worth testing.
Self-service still needs guardrails. A platform should make the routine path easy while escalating tests with unusual legal, safety, marketplace, or statistical risk. Stable feature flags, reusable metrics, approvals, permissions, and automated quality checks let a central team own standards without becoming a ticket queue. GrowthBook's experimentation platform and metric layer support this division of labor.
The central team should act like a product team for internal experimenters. Interview users, measure time to launch and decision, standardize repeated work, and improve error messages. Microsoft's published description of its Experimentation Platform shows the leverage of improving a shared mechanism used across many products.
Reward decision quality, not activity theater
Fabian Hans favors deep dives over mass-produced tests. Once tooling makes launch cheap, raw volume can become a vanity metric. Teams may split minor changes, choose easy-to-move proxy metrics, or avoid uncertain strategic questions. The Cogniteer interview argues that velocity should be judged by useful learning.
Spotify reached a similar conclusion with its Experiments with Learning framework. The framework evaluates whether experiments start with clear intent, use appropriate methods, and affect decisions. A culture that celebrates 500 launches but cannot name what changed is optimizing the visible proxy.
A healthier scorecard includes:
- Percentage of eligible changes evaluated experimentally.
- Median time from question to trustworthy decision.
- Percentage of experiments that pass data-quality checks.
- Percentage that lead to a documented product decision.
- Losses caught before broad rollout.
- New briefs that cite prior evidence.
- Time from decision to flag and code cleanup.
- Cumulative effect on business and user guardrails.
Start with one credible loop
Culture spreads through repeated evidence. Pick one product surface with sufficient traffic, an engaged team, measurable outcomes, and reversible decisions. Define the assignment and exposure contract, create a small governed metric set, run an A/A test, and select a question leadership genuinely cares about.
After the result, publish the full reasoning—including surprises and limitations. Then improve the workflow before inviting the next team. GrowthBook's experiment design guidance provides a practical readiness sequence, while the American Statistical Association's statement on statistical significance explains why context and full reporting must accompany any threshold.
The leaders in this group did not build culture by demanding enthusiasm for testing. They changed incentives, modeled uncertainty, made rigor visible, reduced routine toil, and ensured that learning survived the meeting where it was presented. Do that consistently, and experimentation stops being a specialist service. It becomes how the company decides.
Design for long-term learning
Hear how experienced experimentation leaders balance speed, statistical trust, meaningful metrics, and organizational adoption.
Watch the Leadership SessionRelated Articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics—free.


