Realtor.com on guardrail metrics, the 30/30/30 rule, and AI as your junior data scientist

Guest: Whitney Perez, Director of Product Management, Realtor.com. Host: Ashley Stirrup, CMO, GrowthBook. Show: The Experimentation Edge.
Running an experiment is the easy part. The hard part is knowing whether the result you're looking at is actually telling you the truth.
Whitney Perez has spent her career finding that out the hard way. She started as what she calls a growth hacker at a digital agency, ran A/B testing for the GoToMeeting, GoToWebinar, and GoToTraining sites at Citrix, and helped GoDaddy stand up its first multi-armed bandit for front-of-site promotions and personalization. Today she is Director of Product Management at Realtor.com, where she leads experimentation for the side of the business that serves agents, teams, and other real estate professionals. That team is early in its experimentation program, and Whitney's job is to help it mature.
On this episode of The Experimentation Edge, she walked host Ashley Stirrup through the foundations she's putting in place, the recent test that reminded her why guardrail metrics matter, the 30/30/30 rule she uses to set expectations, and why she believes an English major with a shaky stats record can be a strong experimenter, especially with AI in the loop.
🎧 Listen to the full episode →
Three foundations before you run anything
Whitney's advice for a nascent team starts with getting everyone comfortable with the basics. "What does it take to run a clean and good test? What does a clean and good test look like? What is the value of an A/A test? How do we get instrumented?" Those questions have to be answered before anything more ambitious happens, so that everyone on the team feels empowered by the same shared foundation.
The second foundation is restraint. Not every surface and not every idea belongs in an A/B test. "I am often tempted to, 'I wanna test everything,' and you can't and you shouldn't," she said. Figuring out where the guardrails are, which surfaces are testable and which are not, is a core part of setting up the program.
The third is something she picked up at GoDaddy: a peer review, or bar raiser, program. Every experiment was reviewed by a set of peers. For more complex work, those peers partnered with the data science team. The effect was compounding. People learned from each other, built confidence, and graduated to more complex roles within the experimentation framework. Ashley called it a great example of building culture by getting everyone to help each other, and it doubles as a way to keep the rigor high without bottlenecking every test on one specialist.
The 30/30/30 rule
Part of building that culture is being honest about what the results will look like. Whitney is, in her words, a big believer in the 30/30/30 rule: roughly 30% of experiments win, 30% are inconclusive or insignificant, and 30% lose.
"The math is the math," she said. "That means that the vast majority of your tests are going to be nothing burgers, and that's okay, and that's part of the learnings."
The point of stating the rule out loud is to reset expectations before the first flat result lands. Everyone wants the big splash. But a team that expects every test to win will stop running the tests that teach it the most, and it will be tempted to read significance into noise. Ashley added that the learnings so often come from the losers, and that even a winner is really just a new hill to climb.
The 300% winner that lost money
Whitney's most instructive recent example came from Realtor.com's own checkout flow. The team set up an A/B test introducing a new step for bundling. There were revenue goals attached to bundling, it was a new capability, and, as she put it, there were "all these shiny new squirrels that we wanna chase."
The test appeared to knock it out of the park. A 300% attach rate. Bundling revenue well past target.
Then the team dug in. On the variant with the extra step, attach rate and bundle revenue were high, but the friction of that additional step pushed a meaningful number of people out of the funnel entirely. When they forecast the impact of that fallout against the impact of the attach, the result flipped. "It was actually a loser. The revenue was actually a loser."
Her takeaway is the one experienced practitioners keep relearning: "Yes, your winning metric is really important, but knowing what other impacting metrics and secondary metrics are in the mix and where the thresholds for success around those lies is so important, especially before you're ready to do your victory lap."
Ashley connected the story back to infrastructure. A test like that, read carelessly, could have led the team to believe the exact opposite of what was true. That is the argument for investing in instrumentation and trustworthy data early: the cost of a wrong conclusion is not just a bad feature, it's a bad mental model that shapes the next ten decisions.
Anyone can run a good experiment
Whitney is candid about her own background. "I was an English major, I was a French major. I'm not a stat major. I'm not a math major." Her view is that everyone can be an experimentation champion, and that the path to feeling comfortable in the discipline is shorter than most people think.
Her approach for new experimenters is to strip the definition of a good test down to its simplest form. Understand what makes a good hypothesis. Know what a decision metric looks like. Learn how to avoid confounding variables. Cover the basics of test design: how much traffic you need, what your confidence interval is. Then stop. "Keep it at the simplest definition of a good test and then let them try. You learn by doing in some cases too."
Ashley framed it as the shallow end of the pool being very shallow and the deep end being very deep. Whitney agreed there's something for everyone. Her first experiment was not a multi-armed bandit proof of concept. It was whether a button worked better here or there, and whether the copy looked good. That kind of test can still be impactful, as long as someone on the team brings enough rigor to avoid getting fooled by a result like the bundling test.
Cascading the North Star
Asked how Realtor.com keeps people aligned around North Star metrics, Whitney pointed to something the organization does well independent of experimentation: clarity on the key metrics at the top. The harder work is cascading them down. What lever does an individual surface or team have that moves the North Star? Building that out, whether through an opportunity solution tree or another form of mapping, shows where a movable metric is getting stuck and where the true opportunity lies.
She also emphasized doing the homework before the experiment. Journey maps, usability testing, qualitative research, rage-click data. "Not making the experiment do all the work either. Go do the homework first." Running a test well is a real investment, and pre-experiment design work is how you make sure that investment lands at the right point in the user journey.
AI as an aspirational data scientist
The conversation closed on AI, and Whitney's enthusiasm was specific rather than vague. What she values most is that AI can "turn everyone into, if not a perfect, at least an aspirational data scientist."
She described plugging real data into AI, using the chatbot functionality on the Amplitude dashboard, and asking it to help her think differently about the numbers in front of her, either to analyze results or to surface opportunities she hadn't spotted. The next step is overlaying cross-functional insights, qualitative and quantitative, on top of experiment data to get a richer picture of how one data source relates to the others.
Ashley shared where GrowthBook is focused: giving AI enough context and guardrails that it can help anyone experiment like the best data scientist on the team. He also raised the caution that AI can be very confident while giving the wrong answer, citing Khan Academy teams that used an LLM as a judge and had to be pulled back when the foundations turned out to be wrong.
Whitney's answer was the human in the loop. AI is not going to replace the data scientist or the person reviewing results any time soon, because you still need a sanity check and a smoke test. She drew the parallel to experimentation itself: "You peek in at your experiment after you launch it to make sure it's running the way you expect it to and collecting data the way it's supposed to, and it's the same idea. You can't trust right out of the gate."
Ashley's summary: it's a turbocharger, not autopilot.
What's next for Realtor.com
Whitney's wishlist for her team is three items long. First, perfect instrumentation, because confidence in the data has to come before any experimentation and it's a step you cannot skip no matter how tempting. Second, knowledge embedded on every team, so experimentation is not something only data science or one engineering team knows how to do. Third, a culture where everyone at every level of the organization is excited to hear results, good, bad, and insignificant, because that is what keeps the discipline moving forward.
It's a program built on foundations, honest expectations, guardrails, and a willingness to let people learn by doing. The 300% winner that lost money is the reminder of why each of those pieces matters.
Related articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics — free.




