Realtor.com on using your AI as a junior data scientist
Summary
What do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up.
Chapters
00:00 Cold open and welcome
01:10 From growth hacker to Realtor.com
02:50 Three foundations for a new experimentation team
04:20 The 30/30/30 rule
05:05 The 300% bundling win that lost revenue
07:10 You don't need a stats degree to experiment
09:20 Cascading North Star metrics
12:50 Do the homework before the experiment
13:50 The wishlist: instrumentation, embedded knowledge, culture
15:50 AI as an aspirational data scientist
18:45 Keeping a human in the loop
Notable Quotes
"The 30, 30, 30 rule basically means that 30% of your experiments are gonna be winners, 30% are going to be inconclusive or insignificant, and then 30% are gonna lose. The math is the math."
"When you forecast out the impact of that fallout compared to the impact of the attach, it was actually a loser. The revenue was actually a loser."
"Your winning metric is really important, but knowing what other impacting metrics and secondary metrics are in the mix and where the thresholds for success around those lies is so important, especially before you're ready to do your victory lap."
"Keep it at the simplest definition of a good test and then let them try. You learn by doing."
"Not making the experiment do all the work either. Go do the homework first. Do your journey map, do your usability testing, do your research, and pull in those qualitative insights."
Transcript
The Experimentation Edge - Whitney Perez ===
Whitney Perez: [00:00:00] I'm a big believer in the 30, 30, 30 rule which if you're not familiar with that, it basically means that 30% of your experiments are gonna be winners, which is of course what everybody wants. Everybody is chasing the big win the big splash. But 30% are going to be inconclusive or insignificant, and then 30% are gonna lose.
Ashley Stirrup: Hello, and welcome to today's episode. Today, I'm excited to have Whitney Perez, director of product management at realtor.com. Welcome to the show, Whitney.
Whitney Perez: Thanks, Ashley. Nice to be here, and excited to chat.
Ashley Stirrup: Yeah, you've got a terrific story, so I'm excited to dig in. Maybe we could start off with you just telling us a little bit about yourself and your background, particularly like when you first got into experimentation and things like that.
Whitney Perez: Absolutely. I like to say that I started my career a bit as a growth hacker, and my path to experimentation was very emergent. I came into a digital agency kind of fresh out of college that put on some digital accounts and really liked doing the [00:01:00] merchandising and e-commerce and promotion space.
And at that time, product management and A/B testing weren't as trendy as they are now, but, I slowly built my career around that hopping to increasingly experimentation-focused roles. My first real toe dipped into the water was at Citrix, and I was running the A/B testing on the GoToMeeting, GoToWebinar, and GoToTraining sites.
So working with them to personalize the front of site and make that as personalized as was possible at that time, running A/B tests with the team and optimizing from there. From there I spent a little bit of time in other B2C space and the hospitality space. Did a little bit of experimentation, but the next big milestone in my experimentation career was at GoDaddy.
I helped them set up their first multi-armed bandit and came in to really help run the promotions, merchandising, and personalization on the front of site. That was super [00:02:00] exciting, probably one of the highlights of my career and really I guess earned my spurs so to speak. And then from there, here, bounced around a few more places, and now am at realtor.com, where I find myself at the beginning of the journey with a team that's pretty nascent in their journey towards experimentation and helping them mature and grow and get excited about it too.
Ashley Stirrup: Awesome. And are there any best practices that you're bringing to the team that stand out to you?
Whitney Perez: Yeah, I think there are really three best practices. I think the first is getting clear on the foundations. What does it take to run a clean and good test? What does a clean and good test look like? What is the value of an AA test? How do we get instrumented? Really starting with the basics and making sure everybody is on the same page, and everybody feels comfortable and empowered with those basics before you do anything else.
I think the second is as a best practice, [00:03:00] realizing what does and doesn't make a good AB test. Not everything and not every surface is really suitable for an AB test, no matter how much you might want it to be. I am often tempted to, "I wanna test everything," and you can't and you shouldn't.
And I think figuring out where those guardrails are is probably the second. And third is something I actually picked up at GoDaddy, which is some kind of continuous learning or bar raiser program. At GoDaddy, we set up a peer review program where every experiment was reviewed by a set of peers, and those peers eventually then partnered with the data science team to review if it was more complex so that everybody could learn and build that confidence, and then eventually graduate up to more complex roles within the experimentation framework.
Ashley Stirrup: Yeah, I love that. That's such a great example of helping to build culture by getting everybody else to help each other type of thing. So that's terrific. I think the other thing you did is you kinda helped bring a learning mindset and the expectation [00:04:00] that not everything's gonna be a winner
Whitney Perez: Yeah. I'm a big believer in the 30, 30, 30 rule which if you're not familiar with that, it basically means that 30% of your experiments are gonna be winners, which is of course what everybody wants. Everybody is chasing the big win the big splash. But 30% are going to be inconclusive or insignificant, and then 30% are gonna lose.
And so the math is the math. That means that the vast majority of your tests are going to be nothing burgers, and that's okay, and that's part of the learnings. But I think also having that mindset and knowing that is important going into testing and not expecting everything to win.
Ashley Stirrup: Yeah. Yeah, 'cause so often the learnings come from the losers. So you thought—
Whitney Perez: always.
Ashley Stirrup: Yes. Yeah. Yeah, I'd like to think that the winners you can find a new hill to climb especially if you're getting creative with your tests. But more often the learnings come from the losses.
Can you tell us about an example of an experiment where you had a lot of learnings?
Whitney Perez: Yeah. [00:05:00] There's one that it's pretty recent here at realtor.com, which it was a good reminder for me. We set up an AB test that it was testing out a new step in the checkout flow, introducing bundling. And, we have all these great revenue goals attached to the idea of bundling and what that could mean for us as a business, and it's a new capability and a new, all these shiny new squirrels that we wanna chase.
So we set up the AB test, and it knocked it out of the park. A 300% attach rate , just blew the revenue for the bundling out of the water. But then when we dug into it, we saw that on the variant where that extra step was introduced, although the attach rate and the revenue was very high, it introduced, because the friction of an extra step, a bunch of fallout into the funnel.
And when you forecast out the impact of that fallout compared to the impact of the attach, it was actually a loser. The revenue was actually a loser. And it was such a good reminder that, yes your [00:06:00] winning metric is really important, but knowing what other impacting metrics and secondary metrics are in the mix and where the thresholds for success around those lies is so important, especially before you're ready to do your victory lap.
Ashley Stirrup: Yeah, that's such a great example of, the importance of investing in all the infrastructure you were mentioning before and really having trustworthy data because experiment like that could cause you to think differently about a certain area, where in actuality it's the exact opposite of what you thought you learned if you weren't careful.
So that's a really great example. So in general, how do you help people that are new to experimentation design experiments where they're gonna get the most learnings?
Whitney Perez: Yeah. I say sort of the hard and the soft skills. First, I like to remind people I was an English major, I was a French major. I'm not a stat major. I'm not a math major. This is, not my background and, but everyone can be an experimentation champion. [00:07:00] So I think when people are just starting out, it's helping them realize that there really are easy paths to get comfortable and confident in the experimentation space, and I think AI is accelerating that path, more than ever.
But I think on the hard skill side, helping someone just starting out, it's trying to really get to the simplest definition of a good experiment. Making sure that they understand what makes a good hypothesis, what a decision metric looks like, how to avoid confounding variables, and then how to do some basic test design like, how much traffic do you need?
What's your confidence interval? Those types of things, and leave it at that. Keep it at the simplest definition of a good test and then let them try. You learn by doing in some cases too.
Ashley Stirrup: Yeah. I think in experimentation that's just so important that, you- it's very easy to get overwhelmed or intimidated by A/B testing before you start, and then you do your first one, you realize, [00:08:00] oh, okay.
Whitney Perez: Wasn't that—
Ashley Stirrup: you, yeah, maybe you haven't learned everything, right? 'Cause that, I, that's one of the things I think is so interesting about experimentation is like the shallow end of the pool is very shallow.
It's A or a B or winner or loser. And then the deep end of the pool is very deep, like—
Whitney Perez: Yeah. Yeah. But there's something for everyone. I think that's the really great thing about experimentation. It's not I came in and my first experiment ever was setting up a POC for a multi-armed bandit. It was just, "Hey, does this button, work better here or there? And does the copy look good?"
And it can be that, and that can also be, by the way, super impactful.
Ashley Stirrup: Yeah. Yeah, especially as, yeah, as long as you have the expertise on the team to not get fooled by, an example like the one you just—
Whitney Perez: Yes. Yes,
Ashley Stirrup: yeah. You do still need that rigor step in there somewhere. Yeah. So how do you think about getting people aligned around North Star metrics at realtor.com?
Do, are, is there a lot of clarity around that? And how do people think about, okay, I'm gonna do this thing and maybe it [00:09:00] drives engagement, but my long-term metric is, revenue or something. How do they connect what they're doing to those North Star metrics?
Whitney Perez: I think that's one of the things that we do so well here. Separate even from experimentation and testing. I think that as an organization, we're very clear on the key North Star metrics. And then I think the next step is how do you start cascading those North Star metrics? And this would be true anywhere, but it's certainly true at realtor.com.
But how do you start cascading that North Star metric down to the individual spaces, whether that's surface level, on a platform or at a team level? What lever do you have that you can pull to drive that North Star metric? And then building out, whether it's an opportunity solution tree or, some kind of, mapping where you're you're starting to see what's driving that movable metric and, where it's getting stuck.
Then I think you can start to figure out where the true opportunity lies.
Ashley Stirrup: Yeah. that all really [00:10:00] resonates. A long time ago, I'm forgetting the name of the book now, but it was about, like, how do you get people to do, the new thing while they've still got a day job, and helping them to understand where you're going and then what you're measuring on your way to get there.
Super powerful. It sounds like you're doing a good job of building this experimentation culture within a strong culture already. Are the things you're doing to share wins and losses and, those types of things to just raise the bar in terms of knowledge across the team?
Whitney Perez: It's funny 'cause, so our organization is split into two pieces. So there's the kinda consumer side, so buyers and sellers, and then there's the client side of, agents, teams, and other real estate professionals. And I'm on the side of the house that's, just really starting this journey on the side supporting the real estate professionals.
And so we've had the opportunity rather than me raising the bar, we've had the opportunity to take the baton from a team that is much more mature. And I think people sometimes forget when they're [00:11:00] starting out, there may be learnings, there may be subject matter experts in your org, from whom you can learn and, borrow resources and beg for time to get those learnings and feel like you're not a salmon swimming upstream or having to reinvent the wheel.
And I think that's been, very valuable here.
Ashley Stirrup: Yeah. Boy, that's super interesting. I would imagine that there are aspects of the two sides of the business that are very similar and learnings flow very naturally from one to the other, and there are others where your users and their use cases are just so different that it requires a different mentality.
Whitney Perez: Yeah. I would say even the surfaces and traffic patterns are very different. But I think that's where going back to those basics and saying, what does an experiment look like and where do we experiment? Then you almost get a decision tree almost like a mini, branching yourself.
But you kinda can follow the path to, okay if this is what they're doing over here that might be transferable, then let me walk through those basics and land [00:12:00] myself in a place where it would be applicable or not.
Ashley Stirrup: Yeah. Yeah, that makes a lot of sense. One of the things I feel like having had over 30 guests on the show is that there's some kind of reoccurring themes, and one of them is about really understanding the specific user journey you're working on at that time, and then is the new feature I'm bringing, helping that user on that journey or not?
And that a lot of guests have mentioned coming up with a great idea and then realizing, "Oh, this is the wrong spot," this is still a good feature, but this is the wrong time. They're trying to go here and my feature's trying to take them there, and these are competing. So I'm sure you must have some of that with your business as well.
Whitney Perez: Yeah. And I think even I think of some of that as the almost pre-experiment design work. Not making the experiment do all the work either. Go do the homework first. Do your journey map, do your usability testing, do your research, and, [00:13:00] pull in those qualitative insights. Look where people are rage clicking, use the other tools at your disposal to figure out where the best use of your experiment specifically your A/B test or whatever, is going to have an impact.
Ashley Stirrup: Yeah. Such a great point, 'cause obviously, no matter how much we all try to streamline the experimentation process it's quite an investment to run an experiment well. And so doing your homework, making sure you're investing in the right thing in the right spot in the journey super valuable. Going forward how do you see experimentation evolving for your team?
Whitney Perez: I think that the things that I would love to see on my wishlist is first and foremost, instrumentation. Before you can do any experimentation, the instrumentation and the confidence in that has to be there. You can't skip that step as much as you'd like to as tempting as it may be.
And so I think that's wishlist number one, perfect instrumentation. And then I think wishlist number two item is [00:14:00] knowledge embedded on every team. Sometimes I think that experimentation gets shipped out to only the data science team can do this. Only one engineering team knows how to set up a test.
Only this, only that. And I think every team, every person feeling empowered to run an experiment is the next step in that journey. And then wishlist number three is a culture where, you know, everybody at every level of the org is excited about and hearing the results, both the good, the bad, and the insignificant ones of experimentation to inspire, to help support, to, drive the discipline forward and everything in between.
Ashley Stirrup: Yeah. such a great answer, and I really I'm just super passionate about the whole cultural side of it and how you get people bought in. some of the guests we've had on in the past, they had examples where they've been able to drive a massive amount of revenue [00:15:00] from experimentation.
And once the flywheel gets going, it's very self-reinforcing. But it takes some time and effort to get everybody kinda lined up 'cause especially at more mature organizations, the wins might be small. They might be a half a point here and half a point there, and so each individual win might not be that sexy, but when you stack up a year's worth of wins, you can really move the—
Whitney Perez: looking good.
Ashley Stirrup: Yeah. Yeah. What about AI? Do you see the opportunity to be leveraging AI in new ways?
Whitney Perez: I think there are a lot of cool things and probably a lot that I don't even know yet as the landscape seems to be changing, by the hour at th-this point. But, I think one of the things that I personally love most about AI is that it can turn everyone into a, if not a perfect, at least an aspirational data scientist.
Like I said I like to joke that, I'm an English major. I squeaked by in my stats class. But with AI, I can now plug in, our real data, or I can, use the chatbot [00:16:00] functionality on, the Amplitude dashboard, and I can start to use AI to help me think differently about data and results and the numbers in front of me, either to analyze those numbers or to identify opportunities that, I just didn't spot.
And I think that's one of the most powerful use cases for AI in this space is being able to analyze and look at your data in ways that you just didn't have the hard skills to do before, anybody on any team. And then I think taking that one step farther is overlaying other let's call it cross-functional insights, whether qualitative or quantitative on top of that, and getting a richer picture.
And so being able to say, "Okay, let's look at this experiment or this data source, but how does that relate to these other data sources and other insights?" That's really powerful.
Ashley Stirrup: Yeah. Yeah, I couldn't agree more. At GrowthBook we think a lot about this and I'm sure there'll be multiple [00:17:00] stages to the AI journey, and the one I'm particularly focused on these days is how do you give AI enough context and guardrails so that it can help anybody experiment like your best data scientist?
Over time it'll get better at coming up with new ideas and all that, but some of that will definitely depend on bringing in additional data sources, right? If the AI doesn't have the context, it's not gonna know enough to give you the right recommendations, so—
Whitney Perez: Yeah. And of course, there's a delicate balance there, right? Because, you have to be careful with what data you give it and how you give it, that data and which pieces of that data. And not every company has, even the MCPs to plug in all their data, of course. But I think still a ton of potential there.
Ashley Stirrup: Yes. Yeah, a lot of potential, but, there's also, AI can be very confident as it's telling you the wrong answer.
Whitney Perez: it can. Just, maybe we assume it always has a confidence interval of 80%, right?
Ashley Stirrup: Yeah. But no, I think it's a really valuable point that Khan Academy has done some amazing [00:18:00] work around A/B testing a AI-powered tutor. But they had some teams that were going off and using an LLM and using LLM as a judge, and they had to come in, "Wait. Let's take a look at this.
No, this is completely wrong." And the team felt so confident that they were doing amazing work, but it just wasn't built on the right foundations, so
Whitney Perez: And I think that's where the human in the loop is important. It's never gonna take over the data scientist. It's never gonna take over, the person who's reviewing the results or, at least no time soon probably, because you do need that sanity check and smoke test and it is a lot like experimentation when you think about it.
You peek in at your experiment after you launch it to make sure it's running the way you expect it to and collecting data the way it's supposed to, and it's, the same idea. You can't trust right out of the gate
Ashley Stirrup: Yeah, it's a turbocharger, not... It's not gonna take it o- It's not a, what do they call that when you're driving, just auto steer? It's not auto steer.
Whitney Perez: No
Ashley Stirrup: looking for. It's not autopilot, but yeah. Yeah. Okay. Terrific. Whitney, thank you so much for coming on [00:19:00] the show.
You shared a lot of great learnings. I'm excited to hear where your team progresses. We'll have to have you back on in a year and see how the whole thing has evolved
Whitney Perez: I love it. And I would love it, and thanks for having me. It's always super fun to talk shop.
Ashley Stirrup: Awesome. Thank you
Takeaways from this conversation
Resources
Top takeaways from other favorite conversations

Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea.

Institutionalize learning: align OKRs with “fail forward,” and be willing to kill low-performing features quickly.

Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner.

Start simple if you're new to experimentation; a clean pre/post comparison beats a fancy platform you don't use

Accuracy is a comfortable lie. It grades a narrow test set and can stay high while the agent fails real users.

Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.
.avif)
Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.



.svg.avif)