eCommerce experimentation: Insights and takeaways from the top companies

The strongest ecommerce programs test the economics of the journey, not just the appearance of the page.
Commerce teams can measure conversion quickly, which makes them natural experimenters. It also creates a temptation to optimize whichever click is closest. A winning product-card treatment may reduce basket value. A loyalty discount may increase orders while destroying margin. A cleaner checkout may help new shoppers and confuse returning business customers.
Programs at Fanatics, Box, UPS, Oda, Booking.com, eBay, and toom show how mature ecommerce experimentation connects customer behavior, revenue, margin, operations, and long-term value.
Seven programs and the lesson each contributes
| Company | Product surface | Main lesson |
|---|---|---|
| Fanatics | Sports ecommerce across many sites | Trace revenue wins through diagnostic behavior and replicate surprises. |
| Box | SaaS ecommerce, pricing, cancellation | Choice architecture has boundaries; explain the mechanism. |
| UPS | Digital shipping checkout | Treat operational tools as customer funnels and pair data with research. |
| Oda | Online grocery and logistics | Optimize loyalty and recommendations for profitability, not activity alone. |
| Booking.com | Travel marketplace | Distributed authority needs transparent safeguards and shared memory. |
| eBay | Two-sided marketplace | Measure net marketplace value, not sales shifted between sellers. |
| toom | Home-improvement retail | Make full-traffic product experiments accessible to product teams. |
Fanatics: Verify the revenue story
Fanatics runs close to 100 experiments a month across hundreds of sports sites. Medha Umarji's team does not accept a positive revenue metric without a plausible behavioral chain. When removing ads from product grids appeared to increase revenue, the expected changes in scrolling, product views, and cart behavior did not follow. The team reran the test; the second result was flat.
That is ecommerce rigor in practice. Revenue is a valuable outcome and a noisy metric influenced by large purchases, promotions, traffic mix, and seasonality. Diagnostic measures should show how the treatment changed shopping behavior. The Fanatics program uses repeated tests and meta-analysis to turn individual results into patterns.
Takeaway: Predeclare the path from treatment to revenue. Replicate surprising high-value results when the mechanism does not hold.
Box: Find the boundary of simplification
Danielle Olean's Box ecommerce team found that simplifying a pricing page improved performance, then found that further simplification hurt. The first test did not prove “less information always wins.” It located one better point on a tradeoff between cognitive load and decision confidence.
A cancellation offer also worked in an unexpected way: showing a cheaper plan made some customers perceive greater value in their current plan. The Box examples show why packaging, comparison, and framing need downstream retention and revenue checks.
Baymard's evidence on checkout usability is a strong source of hypotheses about friction. It should inform treatments, not substitute for testing on a company's own products, traffic, and economics.
Takeaway: Convert UX principles into product-specific hypotheses and test the boundary, not only the first direction.
Build an ecommerce scorecard
Choose revenue, conversion, margin, diagnostic, and guardrail metrics that reveal whether a treatment truly improved the journey.
Read the KPI PlaybookUPS: Apply ecommerce logic to operational products
UPS's shipping flow is a transaction journey even if customers do not think of it as a store. Dave Massey's early test removed navigation during checkout and produced an estimated $35 million annualized conversion impact. A later experiment demonstrated that requiring a recipient email could be harmful without context and acceptable when international shippers understood its customs purpose.
The UPS story combines behavioral analytics with UX research. Tests show whether the experience changed an outcome; customer research helps explain the friction and generate the next version.
Takeaway: Map any paid or operational workflow as a customer journey. Test value exchange and comprehension, not merely field count.
Oda: Keep margin inside the experiment
Oda operates online grocery, where discovery, inventory, discounts, fulfillment, and margins interact. The company moved from hard-coded experiments to a warehouse-native workflow on Snowflake and has run hundreds of experiments.
One recommendation improvement increased usage substantially. More importantly, Oda iterated a loyalty program for months before broad rollout to ensure discounts could increase customer value without quietly making the program unprofitable. The Oda customer story shows why contribution economics belong in the decision rule from the beginning.
Google's media experimentation playbook makes a related incrementality distinction: a business should isolate the outcome caused by treatment rather than count activity that would have happened anyway in its experiments guidance.
Takeaway: Pair order and engagement metrics with discount cost, picking and delivery economics, waste, and repeat purchase.
Booking.com: Combine autonomy with visibility
Booking.com became famous for running tens of thousands of experiments a year. The scale comes from distributed authority: teams can launch without routing every idea through a central approval committee. Experiments are visible so peers can question designs, identify collisions, or stop unsafe work.
Its paper on democratizing online controlled experiments links autonomy to technical safeguards, a shared library, transparent data quality, and a central repository of successes and failures. Harvard Business Review's culture account adds the organizational requirement that leaders accept most ideas will not win.
Takeaway: Self-service scales only when assignment, metrics, quality checks, experiment visibility, and stop authority are shared.
eBay: Measure net marketplace effects
eBay's marketplace illustrates cannibalization. A treatment may increase sales for exposed sellers while moving purchases away from control sellers, producing no net marketplace gain. Ordinary user-level analysis can overstate value when inventory and participants interact.
eBay's guidance on A/B test design in a marketplace discusses network effects and why the randomization unit must match the business mechanism. Its research on automated SRM detection also treats assignment validation as production infrastructure.
Takeaway: Define whether the decision concerns buyer value, seller value, gross merchandise volume, contribution, or net marketplace health. Choose the design accordingly.
toom: Expand beyond isolated UI tests
toom wanted experimentation and feature flagging across its retail website without limiting analysis to a small fraction of traffic. Moving to GrowthBook opened full-traffic coverage and created a path for product teams to test more independently.
The toom customer story points toward deeper product questions such as comparing shopping-cart experiences, rather than confining experimentation to cosmetic changes. A warehouse-native approach also lets retail teams use existing transaction and customer data.
Takeaway: Build reusable components and governed metrics so teams can test end-to-end shopping behavior, not only copy and color.
Ecommerce measurement needs special care
Analyze at the randomized unit
If users are assigned, calculate revenue per assigned user. Revenue per purchaser conditions on a post-treatment action and can bias the comparison. Account, household, seller, or market assignment may be more appropriate when experiences spill across sessions or people.
Handle ratio metrics correctly
Average order value, items per order, and revenue per session are ratios whose components may be dependent. Research on ecommerce metrics in online experiments shows that ignoring transaction and item dependence can understate uncertainty.
Cover business cycles
Plan enough runtime for day-of-week behavior, promotion boundaries, payday effects, and data lag. Do not keep a test running indefinitely through major assortment or campaign changes if the underlying environment has changed.
Protect long-term value
Aggressive urgency, defaults, or discounts can improve immediate conversion and damage trust, returns, support, or repeat purchase. Add guardrails and consider a follow-up window or holdout for persistent effects. GrowthBook's holdout framework can estimate cumulative portfolio impact.
Keep operational outcomes visible
Inventory availability, fulfillment time, cancellations, fraud, refund rate, and support load may turn a conversion win into a business loss. Join these facts in the same warehouse-backed scorecard instead of evaluating them after launch.
A practical ecommerce test brief
For the next high-value question, document:
- Customer problem and evidence.
- Journey stage and eligible audience.
- Randomization and exposure unit.
- Treatment difference and expected mechanism.
- Revenue or contribution outcome per randomized unit.
- Diagnostic funnel metrics.
- Customer, marketplace, operational, and trust guardrails.
- Minimum effect worth acting on and planned duration.
- Decision for positive, negative, inconclusive, and invalid results.
- Rollout, rollback, and cleanup owner.
GrowthBook's experimentation platform supports Bayesian, frequentist, and sequential analysis on warehouse data, with reusable fact tables and metrics. Its feature flags carry treatments from controlled exposure to rollout.
The top ecommerce teams do not chase a universal “best practice.” They test a specific value exchange, trace the behavioral mechanism, include margin and operations, and preserve what the company learned. That is how conversion work becomes product strategy.
Keep high velocity trustworthy
Review the power, stopping, SRM, and multiple-comparison practices that prevent ecommerce programs from manufacturing wins.
Read the Prevention PlaybookRelated Articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics—free.


