Experiments

Marketplace experimentation: How DoorDash, Upwork and Expedia run experiments

A graphic of a bar chart with an arrow pointing upward.

In a marketplace, treatment and control often share the same supply. That makes a clean-looking user split capable of producing the wrong answer.

DoorDash, Upwork, and Expedia connect different participants: consumers, couriers, merchants, clients, talent, travelers, properties, and advertisers. A change on one side can alter prices, availability, competition, matching, or behavior on another.

Their public work reveals three complementary lessons. DoorDash publishes advanced designs for network effects. Upwork connects product experimentation to customer feedback, trust, and multi-sided outcomes. Expedia Group has invested in unified analysis and fast circuit breakers that protect revenue during thousands of annual tests.

GrowthBook's guide to treatment effects provides useful language for defining which population effect a marketplace team actually wants, while its experiment design checklist helps surface assignment and metric risks before launch.

DoorDash: Choose the design around interference

DoorDash's dispatch marketplace shares a Dasher fleet across treatment and control. If treated consumers create more orders, they can reduce courier availability for control consumers. Individual outcomes are no longer independent, and a conventional user A/B test can underestimate or overestimate the true market effect.

DoorDash uses switchback experiments that alternate treatment and control across geographic markets and time blocks. This separates shared supply while allowing the same market to serve as its own comparison. The design must handle time trends, carryover, block length, and correlated observations.

The company also balances network effects, behavioral learning, and power. Dashers may need weeks to adapt to a treatment, while longer experiments create greater exposure and environmental drift. DoorDash's discussion of network effects, learning effects, and power makes clear that no design maximizes all three.

Ads introduce budget interference: treatment may consume shared advertiser budget and make control look worse. DoorDash has published a budget-split framework for separating spend and preserving a fair comparison.

Lesson: Draw the interference graph before selecting the randomization unit. If treatment changes shared supply, auction, budget, or matching conditions, individual randomization may answer the wrong question.

Build a marketplace scorecard

Define buyer, supplier, fulfillment, trust, revenue, margin, and diagnostic metrics before competing outcomes reach the results page.

Read the KPI Playbook

Upwork: Combine customer signals with multi-sided product tests

Upwork serves clients seeking work outcomes and professionals seeking opportunity. Its product organization describes customer feedback, research, testing, support, policy, operations, and trust and safety as connected inputs. That structure matters because a change that improves client conversion can increase low-quality proposals, reduce talent success, or shift competition.

Upwork's account of customer-driven innovation says product and customer-experience teams tighten feedback loops and test internally as “Customer Zero.” Public product updates have also reported tests of proposal-boosting mechanics and their relationship to bid and hiring outcomes in an Upwork product update.

The important measurement boundary is match quality. Marketplace funnels do not end at a click, proposal, or contract start. They continue through successful work, repeat relationships, earnings, spend, disputes, and trust. Upwork's work on safe and transparent AI also shows that data provenance, evaluation, privacy, and fairness need to accompany AI matching and assistance.

Lesson: Build paired scorecards for clients and talent, then include platform trust and long-term match outcomes. A one-sided conversion win can reduce marketplace liquidity or quality.

Expedia: Protect the first 24 hours

Expedia Group consolidated experimentation systems from multiple acquisitions into Expedia Group Test and Learn. Its batch analysis supported long-term decisions, but the platform had a gap immediately after launch: harmful treatments could run for hours before a reliable readout.

The company built a near-real-time circuit breaker using Apache Flink. The system aggregates exposure and metric state, handles bot reclassification, and can automatically suspend underperforming treatments. Expedia reports that the real-time monitoring system covered most tests, detected issues within the first day, and stopped a treatment with a severe conversion impact within minutes.

Accuracy matters as much as speed. A noisy circuit breaker that repeatedly kills healthy tests destroys trust and reduces learning. The early monitor also does not replace the batch system; it protects against acute harm until more complete analysis is available.

Expedia's experimentation team has separately published work on sizing ratio metrics, important in travel where revenue per visitor, bookings per search, and nights per booking combine dependent quantities.

Lesson: Treat severe early harm as a production incident. Use conservative automated thresholds, event-time correctness, deduplication, bot handling, and a reliable stop path.

The shared marketplace design process

1. Define all participants and shared resources

List buyers, sellers, service providers, advertisers, inventory, budget, and operational capacity. Mark how treatment can alter availability or behavior for control units.

2. Name the estimand

Decide whether you need the effect on treated buyers, the whole marketplace, suppliers, platform contribution, or a policy at equilibrium. Different questions require different designs.

3. Choose the assignment unit

Options include user, account, supplier, listing, market, geographic cluster, time block, budget partition, or a randomized cluster. The unit should isolate treatment while retaining enough independent observations for inference.

Research on interference in marketplace pricing experiments found materially different estimates under cluster and individual randomization, demonstrating that design can change the apparent effect itself.

4. Model carryover and adaptation

Suppliers learn, reposition inventory, change availability, and adjust bids. Time-block experiments need washout periods or models for carryover. Short tests may capture novelty and miss equilibrium behavior.

5. Use multi-sided metrics

For each side, include participation, success, cost, quality, and retention. Add platform contribution, support, fraud, disputes, and fairness. Decide tradeoffs before launch.

6. Ramp with circuit breakers

Start with a small market or traffic share, validate assignment and exposure, and monitor acute guardrails. GrowthBook's feature flags provide targeting, gradual rollout, and kill-switch control for treatments.

7. Analyze at the independent unit

Do not pretend millions of user events are millions of independent observations when randomization occurred across 20 markets. Use cluster-robust or design-specific analysis and report the effective sample.

Avoid three common marketplace errors

Cannibalization mistaken for growth: A treated seller or ad gains at the expense of control, while total market value remains flat.

Short-term liquidity mistaken for durable value: A subsidy increases matches during treatment but participants leave when it ends or adjust behavior later.

One-sided optimization: Buyer conversion rises while supplier earnings, fulfillment reliability, or trust declines.

DoorDash's research on fractional-factorial business policies shows one path for evaluating interacting policy components more efficiently. The method is useful because marketplace treatments are often bundles, but its assumptions and implementation must match the policy.

Build the data contract

A marketplace experiment record should contain:

  • Assignment unit, cluster, and time block.
  • Eligibility and actual exposure.
  • Treatment version and shared-resource partition.
  • Participant-side metrics and platform economics.
  • Carryover, learning, and washout assumptions.
  • Bot, fraud, and identity rules.
  • Circuit-breaker thresholds and stop events.
  • Analysis method and number of independent units.
  • Result, uncertainty, equilibrium limitations, and decision.

GrowthBook's fact tables can model orders, contracts, searches, assignments, and operational outcomes in the warehouse. Its custom experiment assignment lets teams analyze designs whose assignment happened outside a standard SDK.

The common lesson from DoorDash, Upwork, and Expedia is not a single marketplace test template. It is methodological honesty. The design must reflect who shares supply, whose outcome matters, how behavior adapts, and how quickly the platform can stop harm.

Keep complex tests trustworthy

Review assignment checks, power, stopping rules, multiple comparisons, and causal diagnostics before a marketplace result drives policy.

Read the Prevention Playbook

Table of Contents

Related Articles

See All Articles
Experiments

eCommerce experimentation: Insights and takeaways from the top companies

Aug 17, 2026
x
min read
Experiments

We talked to 4 leaders about getting a stuck experimentation team unstuck — here are their top takeaways

Aug 15, 2026
x
min read
Experiments

We talked to 15 experimentation leaders about losing tests — here are their top takeaways

Aug 14, 2026
x
min read

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.