The Edge Podcast

Battle tested before it reaches the counter: experimentation at Clover

Battle tested before it reaches the counter: experimentation at Clover

Running an experiment is the easy part. The hard part is knowing which experiments are worth running at all, and what to do when the textbook playbook, the classic 50/50 production split, is simply off the table.

That's the situation Ben Schein lives in every day. On this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Ben, Director of Product Management at Clover, one of the largest point-of-sale and payment processing providers in the world. Ben oversees product strategy for the on-premise and digital tools that restaurants run their businesses on: the point of sale a server uses on the floor, the handhelds they carry between tables, and the systems behind menus, promotions, reservations, online ordering, and catering.

The scale is staggering. Clover serves more than 300,000 merchants in the United States alone and processes billions of dollars every day. Ben can watch spikes ripple through the system in real time when a World Cup match or a major concert hits a city. And that scale shapes every decision his teams make about testing.

You can't A/B test a work tool

Here is the constraint that makes Clover's experimentation program different from most: the product is someone's workday.

"Testing is not really 'let's put something in production and see what happens,'" Ben explained. "The worst possible thing is to show up to work the next day and suddenly the tool that you use for work is totally different and no one told you why. You can't do an A/B test in that environment."

A consumer app can quietly ship a variant to 5% of users and watch the metrics. A server picking up a handheld during the dinner rush cannot be a surprise test subject. So Clover inverts the standard model. Instead of testing in production, everything is proven before rollout: structured pilots, ground-level validation, and a detailed go-to-market plan for every feature that ships.

The consequence is that testing carries real financial weight at Clover. "If we don't know that this thing is gonna be successful, we're not gonna invest all those dollars in getting this thing to market," Ben said. Testing is not a stage gate that slows the roadmap down. It is the evidence that justifies the rollout budget in the first place, part of the company's vernacular and culture, as Ben puts it.

Uncertainty and downside set the testing depth

If you can't test everything in production, you have to be ruthless about what earns deep experimentation. Ben's triage comes down to two variables: how much uncertainty is built into the change, and how much real downside exists if it goes wrong.

At the high end sit things like payment authorization flows or the entire funnel a server uses to place an order on a restaurant floor. Those changes are high touch and central to the core experience. They get intentional, ground-level testing before any system-wide change, because when something less proven trips up, the impact is significantly bigger.

At the other end sit table stakes features. Adding Apple Pay matters to the business, but it is an industry standard; the uncertainty was resolved years ago by the rest of the market. You monitor the impacts after rollout, and you move on.

The lesson for experimentation teams: testing depth should follow the risk reward math, not the visibility of the feature. A flashy redesign might need less rigor than an invisible change to an authorization flow. Most roadmaps get this backwards.

Ben applies the same discipline to what happens after a test reads out. His advice to PMs who are new to experimentation is to know the strategic value of a test before running it. "We want the learnings of those tests. We don't always know why we want those learnings," he observed, and a junior PM staring at a significant result with no idea what to do next is the predictable outcome. His fix: structure tests for durable business value rather than pass-fail verdicts, pair every headline metric with counter metrics, and anchor on ground-level measurements like items per check that stay resilient whether $100 or $100 million is being spent on marketing that quarter.

What a burger chain teaches about checkout funnels

Before Clover, Ben led product at Shake Shack, joining at the tail end of the pandemic when the brand's digital channels had gone to market fast and the open question was whether they were any good.

The team's answer became a masterclass in treating a conversion funnel as a brand channel. Shake Shack is built on hospitality, so the question Ben's team asked was strange and productive: what does hospitality look like in a checkout flow?

Some of it was classic optimization: removing redundant steps, and fixing customers who accidentally ordered from the wrong location as store density grew. But the more interesting work was additive. The team tested how presenting prep-time expectations affected conversion, because a Shake Shack burger takes longer than typical fast service and an unexplained wait erodes trust. Set the expectation honestly at checkout, and NPS downstream reflects that the promise was kept.

Then there were the loading screens. Those interstitial moments after a purchase became deliberate brand real estate: the quality of the beef, the fact that the patty is never frozen, conveyed in a passive moment the customer was already spending rather than buried third in a product description. The same thinking carried to the kiosks in the restaurants themselves.

The thread connecting it all is context. "I don't even think of it as personalized, it's contextualized," Ben said. Nobody opens a restaurant app just to browse; they are hungry, they are in the car, they are standing somewhere they can't see the menu. Funnels that acknowledge that intent outperform funnels that don't.

Why this matters for experimentation teams

Ben's career runs from hand-writing test parameters in Firebase config files at NPR, with weeks of work per readout, to a world where the same test runs in an afternoon. His prediction is that the democratization keeps going: interns running hypothetical experiments in sandboxed environments without asking permission, with AI tools making analysis approachable to anyone regardless of background.

But the through line of the conversation is that tooling was never the hard part. The hard part is judgment: knowing which changes carry real downside, structuring tests so their learnings have durable value, and understanding the context your users bring before you measure what they do. Machines may beat humans at allocating ad spend, as Ben readily concedes. Knowing your audience well enough to see where they are going next is still a human edge.

Hear the full conversation with Ben Schein on The Experimentation Edge. And if you're ready to bring that kind of rigor to your own rollouts, visit growthbook.io to see why the open source experimentation platform leader powers testing programs at any scale.

Table of Contents

Related articles

See All Articles
The Edge Podcast
From 22 clicks to 5: the zero impact experiment that shaped how Edd Saunders at JobLeads tests
Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental
The Edge Podcast
Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental
Synthetic audiences meet real A/B tests at Principal Financial Group
The Edge Podcast
Synthetic audiences meet real A/B tests at Principal Financial Group

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.