Filter results

How Fin went from weeks to hours of analysis using AI

Inside The Home Depot's experimentation at a $25B scale

How Disney picks which experiments to run

Ship faster, measure better: experimentation tips from JPMorgan Chase

Twitch on why false negatives kill product ideas

Squarespace killed its blank template and built something better

Signet Jeweler's "View All" page made more money by showing less

RingCentral's DART framework: The four metrics that actually measure AI agents

The 2% close rate increase that turned Ford Credit's product teams into believers

Atlassian on the talent product turnaround from A/B testing

How DoorDash saved millions with one A/B test

How UPS generated half a billion from 80+ Apps with A/B testing

How experimentation led to annual growth at Fanatics
Top takeaways from our favorite conversations

Persistence pays: four months and three to four rounds of trial-model testing at Codecademy produced a 35% conversion increase.

A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.

One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards.

Massey's first test removed navigation from UPS's shipping checkout flow and delivered $35 million in incremental revenue—proving e-commerce best practices apply even when customers think "this is just a tool, not e-commerce."

Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.

Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?

Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.

Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.

Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.

False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.

The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.

Documenting experiments in a centralized Wiki creates a growth flywheel: Fanatics' Wiki feeds their roadmap with iterations on already-built features, reducing tech dependency and accelerating velocity.

Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.

Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.

Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.

Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.

Unblock teams: create a center of excellence for data science and enable rapid variants with AI-powered tooling.





.svg.avif)
