The Edge Podcast

Synthetic audiences meet real A/B tests at Principal Financial Group

Synthetic audiences meet real A/B tests at Principal Financial Group

Running an experiment is the easy part. The hard part is running one at a 150-year-old financial services company, where customer journeys stretch across months, risk tolerance sits near zero, and half the organization believes "we tried that years ago" settles the question.

That is the environment Erika Dunn works in every day. As assistant director of data science at Principal Financial Group, she built the company's experimentation center of excellence and, more recently, something most enterprise teams are still only reading about: synthetic digital audiences, built entirely in house, that just picked a winner in a live A/B test. On this episode of The Experimentation Edge, she walked host Ashley Stirrup through how it all works, and what experimentation leaders at any company can borrow from it.

From observational studies to clean 50/50 splits

Erika's path into experimentation started in quantitative psychology, which gave her the statistical grounding long before she had the corporate title. Her first job studied childhood obesity: observational work with no control groups, just trying to understand what happens when an intervention lands in the real world. At H&R Block, Adobe Analytics brought regular experimentation into the marketing space, and she realized the work she loved in her graduate program had a corporate career attached to it. Then came Amazon, where, as she put it, they can turn anything into an experiment. Eventually she landed at Principal, where she got to dig in and build her own center of excellence around a simple pair of questions: what is experimentation, and how do we share what it teaches us?

That range matters. Someone who has run studies with no control groups and clean web tests with 50,000 users in each arm knows exactly what rigor buys you, and what you lose without it.

When testing gets watered down

The biggest challenge Erika names is not tooling or statistics. It is vocabulary. In more than one company, she has watched testing get watered down as a concept, where "let's try out something" passes for an experiment even though nothing is controlled and nothing is compared. She is careful not to dismiss it: trying things means a team recognizes their ideas need contact with reality. But the difference between trying and testing is a control group, and getting people to appreciate the advantage of rigor takes time, intentionality, and sometimes money. Her core argument is that with just a little pre-planning, you get so much more out of every test.

The second enemy is tribal knowledge. Every organization has it: the insight that lives in one silo and never travels, or the confident claim that "we did that years ago." Ashley recalled a guest from Twitch who spent years being told pricing had already been tested and was off limits, until a rigorous retest showed a huge impact on the business. Erika hears the same pattern inside companies that are 60, 100, or in Principal's case 150 years old. Things change. The segment that ignored an offer three years ago is not today's segment, and a test that lost on the whole population may be a clear winner with high-revenue customers.

Her structural answer is a center of excellence where teams share experimentation wins and losses in one place. Losses are the important part. A visible record of what did not work, and when, is what turns "we tried that" from a conversation ender into a data point with a date on it.

Turning frustration into a metric

Erika's team supports Principal's marketing organization, where the goal is getting the right message to the right people across email and web. The data science advantage, as she describes it, is simple: come in with additional data, or help teams see the data they already have in a new way.

Two examples from the episode stand out. The first is content analysis. Her team broke marketing copy down with readability statistics, including reading level, sentence counts, nouns, verbs, and possessives, and correlated those features with how people responded. That gives content writers something concrete to think about beyond instinct.

The second is the looping metric, developed with her longtime collaborator Josh Ellington, who she is quick to credit. Built in SQL on top of existing web behavior data, it measures how many times a user winds up back at the same main page within tight time frames. When that ratio climbs, it flags desperation behavior: searching, backtracking, revisiting the same section over and over. The team confirmed the signal against session replay tools, and it holds. In tax services, looping predicts bounce. In finance, it predicts something quieter and more expensive: customers who slowly disengage from tools they find frustrating, dragging down engagement and retention over time.

The same discipline shows up in how the team handles assumptions. When they set a nine-second window between an email click and a website visit, it is because Josh dug into the data and found that is how long the trip actually takes. Not five minutes, because a lot can happen in five minutes, and the behavior you attribute across that gap may have nothing to do with your email. Keeping behaviors linked closely in time is unglamorous work, and it is the difference between data and noise.

Customers that look like your customers

The centerpiece of the conversation is Principal's synthetic digital audience program, which Erika describes, for better or worse, as completely Erika built. Drawing on her probabilistic Bayesian background, the system uses real customer data to drive a synthetic data normalization process, producing bell-curve populations of profiles that represent Principal's actual segments without using any real customer data. Each audience is tuned to a specific experimental space: the profiles look like the employer audience in that particular marketing domain, not like a generic average person.

The workflow is deliberately conservative. Before profiles are set loose on new material, they are trained on existing similar content so they replicate known response patterns. Then, through a gen AI process, the team asks each profile its likelihood of opening, clicking, or engaging with a piece of content. Partners can hand over 20 subject lines or banners and get back a ranked list. Teams that want to go wider can generate 50 AI-created variations and let AI duke it out against AI, without burning weeks of human review time.

Erika is precise about what the system does not do. It will not tell you what results to expect. It creates a priority process, narrowing a long list of ideas to the short list worth spending real traffic on. That distinction is what makes the program credible inside a risk-averse company, and it addresses two problems at once. It compresses time in a business where financial journeys are long and feedback is slow, and it creates a safe playground for teams whose tone has always been neutral to finally test new voices, new phrasings, and new content without betting a live campaign on them. As Erika puts it, an experimentation mindset requires that people cannot get punished for making mistakes. They need guardrails, and they need room to learn.

The first win, and three more in August

The credibility test came recently: the first traditional A/B test using a synthetic audience selection ran against a control, and it won. An earlier partner had already seen lift, but without a traditional test design, and Erika's team held their applause until the clean result landed. That partner, working with a small population and a small engagement window, had previously paid outside consultants without moving the needle. The in-house synthetic process is now outperforming that spend.

Three more tests launch in August, after a year spent helping teams across Principal build their audiences. The roadmap points toward agents: Erika sees synthetic audiences moving into agentic workflows, and more broadly, she sees AI helping teams answer the question she hears everywhere now that everyone finally has the data they asked for. We are drowning in information, so where should we experiment next? Her bet is that the next edge in experimentation is efficiency: agents that read your roadmap, your data, and your ideas, and point you at the gains you cannot see from deep in the weeds.

What this means for your experimentation program

Erika's playbook at Principal compresses into four moves any team can apply:

  1. Draw the line between trying and testing. A control group and a little pre-planning turn "let's try something" into an experiment worth learning from.
  2. Turn behaviors you already log into metrics. A looping metric built from existing web data found pain points no heat map had surfaced, and checked assumptions as basic as how long an email click takes.
  3. Use synthetic audiences to prioritize, not replace, live tests. Rank 20 ideas cheaply, then spend real traffic on the strongest, and confirm with a clean A/B lift before celebrating.
  4. Share losses as loudly as wins. A center of excellence with a visible record of failed tests is the only durable cure for "we did that years ago." None of this required buying a platform her company was not ready for. It required statistical rigor, creative metric design, and the patience to build trust one clean win at a time.

Ready to take control of your experimentation program? Listen to the full episode of The Experimentation Edge

Table of Contents

Related articles

See All Articles
Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental
The Edge Podcast
Scaling 100 tests a year across 1,100 dentist owned offices at Aspen Dental
From bottleneck to self serve: scaling experimentation at US Bank
The Edge Podcast
From bottleneck to self serve: scaling experimentation at US Bank
The Edge Podcast
Farfetch's case for building your own experimentation platform

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.