Experimentation
Takeaways
Filter results

The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech.

Add real guardrails: track AI infrastructure costs, ethics/compliance, and inclusion metrics alongside growth KPIs.

Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.

Micro-metrics establish causality beyond top-line KPIs: If revenue moves but scroll depth, cart adds, and product views don't follow the same pattern, question the result before declaring a win.

Start simple if you're new to experimentation; a clean pre/post comparison beats a fancy platform you don't use
.svg.avif)
Measure value by go‑lives and real usage (token volume), not time in portals or playgrounds.

Thumbs-up/down feedback is sparse and skewed. Unhappy users rarely rate — they just quietly stop using the product.

Build a single source of truth (data lake) to power automation and AI reliably.

Use AI call intelligence to score every call against your playbook, surface coaching themes, and save manager time.

Define input and output metrics; ship only what improves core outcomes (retention, sign-ups), and roll back fast if not.

Most B2B product teams are feature factories. The fix is a top-down OKR system, and planning usually breaks in the connections between layers.

A "dry test", a fake "Click here to video chat" button that grayed out on click — measured real demand without building the feature. Of roughly 4 million users, only 106 clicked, killing a multimillion dollar build.

Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs.

The most valuable experiment work happens before you push play: clear enrollment logic, a plain-English hypothesis, and no optimizing ahead of the test.

Synthetic digital audiences rank 20 content options by predicted engagement before a single live impression; Principal's first selection beat the control.

The real bottleneck is alignment, not developer resources. Agree on the problem and its hierarchy before anyone builds a variation.

Exploration is part of the customer's delight; returning customers wanted to browse the menu even though they ordered the same thing every week.

One or two big wins a quarter is a healthy hit rate when you run 150–200 experiments a year.

Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.

Institutional memory is infrastructure — every test result since 2020 lives in a centralized, searchable archive so no one re-runs a question the company already answered.

Metrics should be driven by the experiment's hypothesis, not chosen by leadership in a silo. Pair a primary KPI with secondary KPIs for return behavior.

Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut.

Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test.

Experimentation short-circuits political debates by removing opinion from product decisions.

Reframe experiment outcomes as savings and gains, not wins and losses. A "losing" test saves you from a costly mistake, which keeps teams focused on learning instead of fearing failure.

Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged.

Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis.

In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams.

When you can't test at scale, desk rides replace A/B tests — sitting with users and watching them struggle reveals failures faster than any dashboard.

False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.

Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top.

Fin A/B tests everything, even one-character prompt changes, and treats a 20 to 30 percent win rate as a sign of a healthy program.

Ship first, then optimize: launch PLG features and immediately run experiments to increase adoption; track daily active usage per feature.

Stated preference lies: users asked for a blank canvas, but behavior demanded guided design — and only the experiment could referee.

Enforce experiment hygiene: change one variable at a time, randomize at the right unit (account vs. user), and run long enough for effect size.

Shift quality left with automated checks so developers catch issues early without human gatekeeping.

An experimentation mindset requires that people can't get punished for mistakes; guardrails plus a safe playground beat running the same test forever.

Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy.

AI scales institutional knowledge, not just analysis speed — mining past experiment readouts to auto-generate new hypotheses turns your testing history into a compounding advantage.

A simple fallback, like a two second load rule, can save an ambitious experiment without sacrificing coverage or security.

What you cannot measure, you cannot ship — if you can't measure an outcome, you can't decide whether it's better, so you're just debating opinions.

Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume.

Keep humans in the loop for AI-assisted coding and customer answers—trust but verify in regulated contexts.

More conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations.

Build hypotheses around user psychology, not just KPI movement

The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.

Accept that being wrong is the point—experimentation only works when leadership embraces humility

UPS runs everything centrally now, but the real win is that demand for testing has decentralized—business units across the company now come to J.E.D.I. asking to test their ideas.
.avif)
Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.
.avif)
Fewer, bigger experiments beat high volume. Signet went from 40–50 tests a quarter to 15–25 because complex, value-driven tests produce reusable insights that small tweaks don't.

A looping metric built from web data finds where customers get stuck without heat-mapping tools: watch how often users cycle back to the same page.

Accuracy is a comfortable lie. It grades a narrow test set and can stay high while the agent fails real users.

A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new.

Decide testing rigor with blast radius x reversibility; reserve heavy testing for irreversible, high-impact systems.

In a regulated industry, every customer must be accounted for. Even one to two percent of users missing an experience is unacceptable.
.svg.avif)
In low-volume B2B, read losing experiments for sub-segment signal; a "failed" Stripe form simplification revealed the form was blocking legitimate small-business buyers using Gmail.
.svg.avif)
Evolve evals for agents: track tool call accuracy/success and task completion/adherence; A/B test models and strategies.

Run a broad explore experiment first; small, over-narrowed populations lack power and raise the odds of a false negative. Find the responsive segment with heterogeneous treatment effects afterward.

If an intervention sounds weak when you write it out in plain English, don't run the experiment — you're just wasting time.
.avif)
Tie every result to dollars. Translating experiment outcomes into revenue is how Craig keeps financing, warranty, and chat stakeholders aligned and gets executives to act.

AI is ushering in a golden era for experimentation, because shipping faster only compounds mistakes unless you measure what you ship.

Moving new experimenters from solution space to problem space thinking raises win rates and produces learnings the whole organization can use.

Route every experiment through one entry point. Farfetch's feature toggle connects segmentation, user systems, CMS and messaging.

Turn data into narratives with AI to deepen engagement and increase discovery.

Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.

Make experimentation company-wide: centralize data (BigQuery), broadcast wins/losses in Slack via GrowthBook, and auto-correlate metric dips to releases.

Start with low-risk, high-yield AI use cases—unit tests, documentation, and security triage—to build confidence and momentum.

A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control.

AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company.

Use AI to accelerate builds, detect incidents sooner, and evaluate models; watch MTTR and MTTD.

A center of excellence should enable, not execute. Farfetch's central team shrank while experiment volume grew, because its job is coaching.

DART measures behavior, not opinions. Four signals read off logs and transcripts: decay, acceptance, relevance, and task completion.

The same feature (required recipient email) failed for customer data capture but passed for international customs—proof that framing and customer benefit matter more than the feature itself.
.svg.avif)
AI has collapsed marketing analysis from weeks to hours, and the real payoff is a cleared experiment backlog plus analysts who compete on the questions they ask, not the speed they query.

Scaling experimentation from 0.3 to 2.8 tests per month is less about education and more about habit change, shared learnings, and giving non specialists the tools to launch their own experiments.

Cut time-to-lead with workflow automation and track the downstream impact on conversion.

Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.

Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner.

Product to channel fit decides what sells online. Books and fashion judge well on a screen; perfume and washing machines need cues a screen cannot give.

Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.

Build an AI ecosystem with clear purposes (productivity, engineering, consumer) and a steering committee to avoid duplication.
.svg.avif)
Productize adaptability with model routing to match tasks to the right model family as capabilities shift.

Test “obvious” UX changes; preserve helpful friction and align with user mental models.

A control group is non-negotiable: at scale, a change worth millions is invisible under noise and seasonality, and no one can spot it by eye.

Strategic bets deserve a longer clock than fail fast allows. Farfetch iterated on Inspire for two years before it replaced the market leader.

Scale test volume to learning speed, not just shipping speed

Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways.

Getting a stuck team unstuck starts with data and a workshop. A Disney team went from "we don't know where to start" to 110 scored, prioritized test ideas, using Contentsquare heatmaps to diagnose low engagement first.

Treat engagement carefully. For a bank, more time in the app isn't a win; trust, fast task completion, and healthy repeat engagement are.

Frameworks like "do no harm" and "small sample" expand who can test: Not every initiative needs 30,000 orders to ship value—lower the barrier for teams that can't hit statistical thresholds while protecting core KPIs.
.svg.avif)
Build self‑verification into workflows: pair agents with automated testing (e.g., browser runners) and iterate to thresholds, not perfection.

Measure DORA metrics and developer sentiment; remove mundane toil to increase speed and satisfaction.

Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.

Separate your two experimentation modes: high-volume CRO chases many small wins, while big uncertain bets deserve multiple shots to de-risk.

Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?

Documenting experiments in a centralized Wiki creates a growth flywheel: Fanatics' Wiki feeds their roadmap with iterations on already-built features, reducing tech dependency and accelerating velocity.

Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.

Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.
