The Edge Podcast

Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental

Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental

Running 100 tests a year is the easy part. The hard part is knowing what a win actually means when the "customer" on the other side of your website is 1,100 different businesses.

That is the problem Arie Polycarpou works on every day as Senior Manager of Test and Learn at Aspen Dental. On this episode of The Experimentation Edge, he walked host Ashley Stirrup through twelve years of building experimentation programs at Kohl's, Marriott, Total Wine, and now the largest company in The Aspen Group's healthcare portfolio. Along the way he shared the two numbers he actually manages to: a win rate he deliberately expects to fall, and stacked win math with a haircut built in.

twelve years from volunteer to program builder

Arie's route into experimentation started with a hand raised at Kohl's Department Stores. He had landed in digital analytics out of a marketing analytics background, drawn to digital because, unlike most data work, you get to see a whole customer experience start to finish. A/B testing was a growing field at Kohl's with exactly one person working on it. That person needed help. Arie volunteered.

Twelve years later, the pattern of his career is remarkably consistent: join as an individual contributor, connect the data sources, install the rigor and the processes, then build the team. At Marriott he helped bring measurement discipline to A/B tests. At Total Wine, a company that grew tremendously during COVID, he took a program doing a couple of big product tests a quarter and turned it into a full-fledged, hundred-test-a-year operation. After graduate studies in digital transformation, he joined Aspen Dental in October 2025 to bring its agency-run testing in-house and accelerate both the volume and the depth of what the program can measure.

Now he leads a team of three: himself, a manager of analytics, and a coordinator of testing and analytics. The descriptive work matters, but the bread and butter is turning known business opportunities into tests, then working with product managers, the merchandising and production group, and a team of front-end developers to build them. The target: about 100 tests a year, with room to grow.

healthcare retail: one website, 1,100 offices

Aspen Dental is not a typical B2C business. It is part of TAG, The Aspen Group, which owns several healthcare organizations, and Arie describes the model as healthcare retail. Aspen Dental has about 1,100 offices across the country, and none of them are corporate-owned. Private dentists own the practices, in a model Arie compares to Marriott: the brand provides resources, training, and education, but the practice owners run the businesses.

The website's job is focused: attract new patients, help them learn about the offering, and get them on the books. Appointment bookings per location is the North Star. As Arie puts it, the website's role is to "give the offices a chance to serve a user."

But that franchise structure changes what a win means. Every office has its own star rating, level of service, show rate, and patient return value. A test that lifts bookings overall can still be a bad trade if it routes more patients toward weaker offices or lower-value visit types. So Arie's team monitors what happens after the booking: the value of a patient, whether the mix of users is changing, and whether the mix of offices receiving those patients is changing. Promoting stronger offices, with higher show rates and higher return value, is simply better for the business than spreading demand indiscriminately.

For experimentation teams, the lesson generalizes: when your business is really a network of businesses, the unit of analysis has to follow the value, not just the conversion.

the win rate that should fall as you mature

Ashley asked Arie a question every experimentation leader gets from their executives: what's your win rate?

His answer runs against instinct. Arie strives for a true win rate of about 25%, and he treats the roughly 40% win rate that new programs often post as a symptom of youth, not excellence. Early programs feast on low-hanging fruit: the obvious fixes, the long-postponed ideas, the big value opportunities. As those get knocked out, the tests get more ambitious and the win rate should decline.

The logic cuts both ways. Too low a win rate means you're wasting effort testing the wrong things. Too high means you're not testing anything hard. And the baseline is stiffer than a coin flip suggests: your current experience already embodies years of descriptive data, UX best practices, and accumulated knowledge. Beating it significantly takes real work.

The industry's most mature programs bear this out. Ashley noted that Bing has been quoted around a 20% win rate, and Airbnb as low as 10%. The more mature the product, the harder it is to move the needle. A falling win rate, in other words, can be evidence that your program is finally testing things worth testing.

the haircut rule: two 5% wins don't make 10%

The second number Arie manages carefully is the one that gets reported upward: the stacked, annualized value of the program's wins.

His rule is to be optimistic about better and conservative about how much. A 5% lift measured during a test window rarely survives a year in market. Customers acclimate and the novelty gets baked in. New changes get layered on top of old ones. And lifts don't compound the way arithmetic implies: a 5% homepage win followed by another 5% win does not produce 10%. "I know the math should make it seem like it does work like that," Arie says, "but it doesn't."

So depreciation gets baked into estimates before they reach leadership, in the form of haircuts. The approach echoes what Ronny Kohavi described on a GrowthBook webinar: apply a 20% haircut when you stack wins. Arie goes further when the evidence is thinner, taking a bigger haircut on tests that ran shorter or won with weaker confidence.

The alternative, long-term holdouts, is statistically cleaner but operationally messy. Holdouts force you to maintain two experiences, effectively two sets of code, and the longer they run, the harder they get. For most programs, disciplined conservative estimation is the more sustainable path, and it pays a cultural dividend: conservative numbers that hold up build trust with stakeholders in a way inflated dashboards never do.

where the program goes next

Aspen Dental's program is still young; Arie joined less than a year ago and the muscle is still being built. The road map runs toward specificity: localization, where each of 1,100 offices can be served differently, and service lines, where the experience that books a denture repair is very different from the one that serves a dental emergency.

On AI, Arie is interested but deliberate. Testing tools are adding AI to analyze tests faster, look up results faster, and even help build tests. He sees the opportunity, and Aspen is looking into it, but with guardrails: keep a clean code base, keep top engineers in control of the site, and use AI as additional analytics hands that free the team to think more strategically, not as a replacement for rigor.

why this matters for experimentation teams

Three disciplines run through this conversation. Follow the value past the conversion event, even when that means office-by-office analysis. Read your win rate as a diagnostic of maturity, not a scoreboard. And report stacked wins with a haircut, so the number you promise is one that holds.

Ready to take control of your experimentation program? Start for free or get a demo at growthbook.io.

Listen to the full episode of The Experimentation Edge

Table of Contents

Related articles

See All Articles
Synthetic audiences meet real A/B tests at Principal Financial Group
The Edge Podcast
Synthetic audiences meet real A/B tests at Principal Financial Group
From bottleneck to self serve: scaling experimentation at US Bank
The Edge Podcast
From bottleneck to self serve: scaling experimentation at US Bank
The Edge Podcast
Farfetch's case for building your own experimentation platform

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.