The Edge Podcast

The four questions Early Warning asks before any A/B test

The four questions Early Warning asks before any A/B test

Running an experiment is the easy part. The hard part is knowing whether it was worth running at all.

Priya Singhee has spent a decade finding out where that line sits. As global head of storefront product analytics at Wayfair, her teams tested every step of the funnel, from the moment a shopper searched for a couch on Google to the moment they checked out. Today she is VP of Enterprise Analytics & Data Science at Early Warning, the 30-year-old consortium owned by the seven largest US banks that fights identity and payment fraud, and the operator of Zelle, which processed a trillion dollars in payments last year.

On The Experimentation Edge, she told host Ashley Stirrup something most experimentation teams don't want to hear: the majority of testing dysfunction happens before the test ever launches.

Listen to the episode here.

The framework: can you test this validly?

Singhee's decision framework starts with a reframe. The question is never whether a change deserves a test. "It's really not about should we test, but it's more about can you test this validly?"

Four questions decide it.

First, can you truly randomize? "It truly only works when the randomization is clean," she says — clean separation between who gets treatment and who was going to succeed anyway. If cross-pollination, unrandomized seeds, or a mismatched randomization unit contaminate the split, there is no case for a standard A/B test. The unit question is subtler than it looks: teams routinely randomize on one unit — session, user, device — and then read out results on another. That mismatch alone invalidates a readout.

Second, is your effect size plausible given your traffic? "If you only get 5,000 checkout sessions a month, you're hoping for a half percent lift in conversion — you'll need months of data." Most people's guesses about detectable effects are wrong, and the power math has to happen before launch, not as an afterthought.

Third, is the change reversible and cheap to test? A return-policy change, for example, contaminates so much of the experience that a clean split becomes nearly impossible. Some changes simply aren't testing candidates.

Fourth, and the one Singhee weighs most heavily: do you actually have a hypothesis, or are you just poking around? "A test without a specific falsifiable hypothesis is just like a fishing expedition in my mind." Writing it down — if we cut checkout from four steps to two, conversion improves by 1.5 points — forces the team to articulate a mechanism and creates the standard the result gets checked against later. It is also, she notes, the single best vaccine against p-hacking.

The statistic nobody builds their culture around

Here is the number that should reshape how every experimentation program measures itself: 85 to 90% of tests fail. It's a well-known statistic. Yet in most companies, the pressure runs entirely the other way — hunt for wins, stack up wins, report wins.

Singhee flips the incentive. "You have to understand the A/B test is a learning agenda. If you've learned something, it's good enough." And the corollary cuts deeper: "If you're winning too many, I would be very skeptical, because really 85 to 90% of them are supposed to fail." A suspiciously high win rate isn't evidence of a brilliant team. It's usually evidence of a broken measurement pipeline — novelty effects read too early, hidden segment effects miscalled as wins, or multiple comparisons quietly inflating false positives.

Her prescription is to decide the endings before the story starts. Every pre-registered analysis plan should answer three questions: What will we do if it succeeds? What will we do if it fails? What will we do if the primary metric wins but guardrail metrics decline? Getting VP-level approval on those answers up front means nobody is left "desperately trying to prove it to be a win" after the fact.

The best organizations she has seen go one step further: they log every test and every learning in a place the whole company can see. "That is a world-class organization where there's this reinforcement loop going on... And that's how you learn as an organization."

Pre-registration is the antidote to p-hacking

P-hacking rarely announces itself. In Singhee's experience, it arrives politely, after a test has gone sideways, in the form of requests to the analytics team: Can you look at just new users? Just the mobile segment? Just this one slice?

"You will start with your power and duration on a particular sample and a detectable size," she explains, "but nevertheless, you'll then start wondering, 'Oh, how would it do on new users? How would it do on just the mobile traffic?' You never pre-registered for that, so you can't be reading out your experiment results on those."

The fix has to be installed before launch. Her analyst checklist: the hypothesis, the mechanism, the primary metric, the exact statistical test that will be run, and every subgroup analysis the team plans to perform. Anything outside that list is exploration, not evidence.

Then comes the discipline most teams skip entirely: a pre-registered kill criteria. "What would make you say out loud that this idea was wrong, we're not shipping it?" Without a written answer, the default human behavior is to hunt for a subgroup where the losing idea secretly worked.

None of this is about winning more. "The biggest thing for me in a learning agenda is what did you learn from this test? I don't care if it won or lost."

Losses are where the money is

Singhee closed the conversation with the argument she wishes every skeptical executive would internalize. If 85 to 90% of ideas fail, then a company that ships everything without testing is silently absorbing all of those losses. "Imagine if you didn't A/B test... you'd actually be losing revenue." The wins get the headlines, but loss avoidance pays the bills: ship ten features untested and the two losers can erase everything the two winners gained.

Her sincere request: "People do more A/B testing, not less."

The teams that get this right won't be the ones with the highest win rates. They'll be the ones who can answer, for every test they've ever run, one simple question: what did we learn?

Ready to put a real decision framework behind your experimentation program? Start for free or get a demo at growthbook.io.

Table of Contents

Related articles

See All Articles
The Edge Podcast
A/B testing 300 million players without breaking their trust: the Supercell approach
Battle tested before it reaches the counter: experimentation at Clover
The Edge Podcast
Battle tested before it reaches the counter: experimentation at Clover
The Edge Podcast
From 22 clicks to 5: the zero impact experiment that shaped how Edd Saunders at JobLeads tests

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.