Back to Podcast
A/B Testing
ROI
Scale

PayPal's $180 million experimentation win

S1 | E46
Sep 30, 2026

Summary

Gaurav Sethi joins Ashley Stirrup on The Experimentation Edge to explain how PayPal's experimentation program went from 700 to 800 tests a year with four week readouts to 2,587 experiments a year and roughly $180 million in measured impact. Gaurav inherited a reported win rate of 55% to 60%, four times the industry average, and traced it to experiments logging assignment data instead of exposure data, carrying 25% to 30% dilution. The conversation covers the exposure event his team introduced, the instrumentation and metric standards that cut readouts to 24 hours, a carousel test that a multi armed bandit resolved in 51 days instead of a projected 700, and why cost avoidance from losing experiments belongs in the ROI number. It closes on what changes when the thing you are testing is an AI agent rather than a button. Useful for product managers, engineers, data scientists and growth leaders building or defending an experimentation program at scale.

📝 Read the full blog post →

Chapters

00:00 Cold open
00:54 Welcome Gaurav Sethi of PayPal
01:30 Elmo, PayPal's homegrown experimentation platform
03:53 $180 million in revenue impact and cost avoidance
04:28 The win rate that was too good to be true
06:29 Instrumentation standards and the exposure event
08:56 The carousel test: 700 days down to 51
11:29 Designing experiments so every result teaches you
13:45 Building the platform is only half the job
15:07 Education, office hours and executive support
18:28 Exposure events and joining transactional data
24:14 AI for experimentation, and experimentation for AI

Notable Quotes

"If you go to paypal.com, as soon as you land on the page, you will be assigned a control or a treatment. That is the assignment or the intent that we are going to show you control or treatment. But that does not mean you actually saw the control or treatment."

"If the data is diluted, then you will reach stat sig very quickly, even if some of the users did not even participate in that experiment. That was one of the big reasons it was creating a false win rate for us."

"There are no winning or losing experiments. If your experiment is winning, that means you're going to have a direct revenue impact. But if your experiment is losing, that also is sort of a winner because that is helping you to not ship a bad piece of feature. One is a direct impact, one is a cost avoidance."

"Building an experimentation platform is just half of the job. If we put a good set of engineers together, maybe in six months you can build a decent enough experimentation platform. What took a lot of effort was educating teams on the best practices of experiment."

"I think by default, every agent that should go live in production should go behind a feature flag. If an agent has access to a lot of data and a lot of context and it is not governed, that could actually be pretty dangerous for some of the businesses."

Transcript

Ashley Stirrup: Hello and welcome to today's episode. Today I'm excited to have Gaurav Sethi on the show. Gaurav has spent over eight years at PayPal and before that he was with Hitachi Data Systems, rebranded as Hitachi Vantara. So welcome to the show, Gaurav.

Gaurav Sethi: Thank you Ashley, thank you for inviting me to your episode.

Ashley Stirrup: Yeah. great to have you on. to kick things off, maybe we start with you telling us a little bit about your experience at PayPal and what experimentation was like when you got there and how things evolved.

Gaurav Sethi: Absolutely. So first of all, PayPal has a homegrown experimentation platform. We named it as Elmo, referred to as Experiment Lifecycle Management and Optimization Platform. And as the name suggests, it allows you to manage the entire life cycle of the experiment. I inherited this team earlier in 2023. At that time, we were running probably 700 to 800 experiments a year. A lot of things were not going well because one of the biggest challenge we saw was insights. The experiment insights were taking way longer like sometimes up to four weeks because a lot of the data methodologies we were using, the data collection we were using was kind of broken and that was taking data scientists or data engineers together to clean up all the data first and then finally being able to get to the insights part of that. And the other missing part of that platform was at that time There was no financial metrics associated with an experiment. There was no way to marry the transactional and the behavioral or the experimental data that we were collecting to be able to do that. Last year, just to fast forward last year, we ran about 2,587 experiments, to be precise. Of those, 55 % to 60 % were ABA experiments. a lot of them, other than from the rest of the 45%, some of those were disguised as a feature flag. So a lot of teams were kind of abusing this as well as a feature flag. And then we had 1 % or you can say 2 of experiments running like hierarchical experiment, which is something pretty unique to our platform, which we built to help developers, to make life easy for developers. And lastly, a couple of MAB tests and some new features that we were testing was the network experiment effects and all. By the time I left PayPal, we were running about, the experiment reroutes were available within 24 hours. I had already built a layer where you could define the metrics for an experiment. All of the automated jobs would collect all the data, marry that to the transactional data, and give you a financial impact as well. And we will dive into these, I hope, in the rest of the episode.

Ashley Stirrup: That sounds great. Yeah. Sounds like you had a massive impact on the business as well.

Gaurav Sethi: It was, yes. In 2025, the revenue impact we created was about $180 million. And it can be divided into two pieces. One is the actual revenue impact that we are making with all the winning experiments, and also the cost avoidance that we achieved by not shipping any features or any experiments that were counted as a failed experiment or a not losing experiment. So combining both of these together, was somewhere 180 million dollars.

Ashley Stirrup: just enormous. And and what was the typical win rate there?

Gaurav Sethi: That's an interesting question. So I think in 2024, data scientists were claiming that we had a win rate of about 55 % to 60%. That something was not really, I would say, not correct because even industry average for that is about 11 to 14%. So 55 % win rate was something that did not really sit well with me. When I went in and kind of tried to... look under the hoods what was going on, I identified that a lot of those experiments were using default tracking methodology. As you know, there is an assignment data and there is an exposure data. So a lot of those experiments were using assignment data. And when I dug further, I found that most of those experiments using assignment data, they had about 25 to 30 % of dilution as well. And just for the listeners, just to kind of get into that, look, if you go to paypal.com, as soon as you land on the page, you will be assigned a control or a treatment. That is the assignment or the intent that we are going to show you control or treatment. But that does not mean you actually saw the control or treatment. You never probably navigated

Ashley Stirrup: Got it.

Gaurav Sethi: to that section of the experiment. You never probably went to the mobile page where the experiment was running. And that's what was diluting our numbers. So. Later on, when I fixed that, I started making sure that we started educating teams on using the exposure data for analysis. This win rate came down as well.

Ashley Stirrup: That's that's super interesting. Why would like a diluted exposure rate lead to more wins? Did it just make you overconfident on whether you had stat SIG or not?

Gaurav Sethi: Correct, because if the data is diluted, then you will reach static very quickly, even if some of the users did not even participate in that experiment. So that was one of the big reasons it was creating a false win rate for us as well.

Ashley Stirrup: Got it, got it. And how did you address the data issues? It sounded like that's a pretty meaty problem for you.

Gaurav Sethi: Yes, it was. So there were actually a bunch of things that were going against the data quality that were going wrong, you can say. And it all starts with instrumentation because it was very difficult for all the engineers to align on a specification of what they want to instrument. Like checkout would have a different way to instrument. They had different metrics to track. They had their own internal metrics that they were tracking, their leaders were tracking. Whereas crypt and B2P or a digital wallet would have different metrics that they are tracking. So aligning everybody on a standardized metric and on a standardized instrumentation specification was one of the very big challenge. Now, since we were also trying to automate all of this process, whenever we saw the instrumentation specifications were not correct, that meant that a machine or an automated pipeline that we had that were running all these analysis cannot really function without human intervention.

Ashley Stirrup: Yeah.

Gaurav Sethi: where all the data scientists or data engineers had to intervene to be able to get the right insights. So first of all, aligning on that instrumentation specification, then aligning on assignment versus exposure. I had to educate a lot of team members on what should be a trigger point for your experiment. And this is where I also introduced the concept of an exposure event. What that meant was as close as you are to render the experiment to a user, fire an event that will tell us that from this point onward user is part of the experiment. Don't fire the event or don't log the data as soon as user lands onto the page.

Ashley Stirrup: Yep. Yep. Makes a lot of sense. Sounds like an important part of that was kind of creating a common, you know, view on how you track information, getting everybody to buy into doing it the same way.

Gaurav Sethi: That's correct. And I think that was one of the challenging part as well, because as you can imagine, checkouts number one priority is not logging the right data for experimentation. Their priority is building the features that checkout is, the strategic features that checkout is planning to do. So aligning with them on certain standards, aligning on what exactly you should be logging, when you should be logging, and making it easier for them as well via SDKs, via exposure event, any of that. All of that was built in into the platform also and as part of the SDK as well.

Ashley Stirrup: Makes a lot of sense. Do you have an example of an experiment you ran where you had a lot of learnings?

Gaurav Sethi: think the MAB experiment that we ran was one of the most fascinating experiments we saw. So on PayPal.com home page, we were having a carousel with images and text both. And marketing team wanted to try different options by changing the image as well as the text. And that created about six or seven variations that they wanted to test. When we looked at the stat sig, like how much time it will take to get to stat sig, looked like about 700 days will take for us to get stats to validate all of these six different variations. And that is where we came up with an idea to run an M-A-B experiment and multi-arm bandage.

Ashley Stirrup: Yeah, multi multi yeah, multi arm bandit. Yeah. Yeah.

Gaurav Sethi: The way it works is, the way I like to call it is, it's an A-B experiment on steroids. So what it does is, when you have a bad performing arm or a bad performing variable, variant, it takes the traffic from the bad performing variant and distributes among the rest of the variants. So it's like an explore and an exploit methodology that is used. When we launched that, we identified that within 40 days or so, we were able to find an interim winner. And by interim winner, I mean that there were two or three variants against control. There were two variants that were performing way better than everybody else. So we could eliminate that. right away. And next two weeks or so we ran A-B test on just those two variants to be able to see which one is winning faster. And by I think day 51 we were able to get the static results, we were able to identify the winner that had almost five basis points of improvement as compared to everybody else.

Ashley Stirrup: Wow, that's huge. So you would think PayPal.com gets a lot of traffic and so it wouldn't take seven hundred days to get to stat sig. do you you have any color on why it would take that long?

Gaurav Sethi: Yes, because a lot if you think about it, when users land on like, first of all, users started using mobile more recently. So not everybody was going directly to paypal.com to be able to do that. Second was

Ashley Stirrup: Mm-hmm. Right.

Gaurav Sethi: a lot of the time, the way we get traffic is when you say pay with PayPal. So you land directly onto a login page, you don't really land onto the homepage or the hero banner that we have. So in that

Ashley Stirrup: Got it.

Gaurav Sethi: sense, the traffic that we were receiving was way smaller than it actually seems for rest of the PayPal traffic.

Ashley Stirrup: Yep, that makes a ton of sense. And let's say you were gonna work with a new product manager, they're all excited about their feature, they're ready to test it, they're just sure it's gonna be a winner. how do you help them make sure they're designing the experiment so that if it's a loser, they can learn as much as possible about it.

Gaurav Sethi: Yeah, I think so. Over a period like when I started learning about experimentation as well, I finally understood that there are no winning or losing experiments because if you're winning and if your experiment is winning, that means you're going to have a direct revenue impact or direct KPI, whatever you're tracking, conversion rate, click through rate or any of that. But if your experiment is losing, that also is sort of a winner because that is helping you to not ship a bad piece of feature or a feature that you think is not going to ultimately scale for you. So in that sense, both of them are winning experiments. One is a direct impact, one is a cost avoidance. How do I help my product managers is, first of all, I make sure that they understand the basic guidelines of running an experiment. To give you an example, I have seen so many of the experiments that were set up with one-liner hypothesis, testing a red color button just to see if user clicked more. And that did not show anything. It's very vague in many sense. So educating them on how you should be writing a hypothesis. Then also helping them design what should be your target segment that you're looking for. How should you be identifying to make sure that there are no biases on your segment? And finally, making sure that you have a goal associated before even you launch an experiment. And by goal, I mean like a primary KPI. That will be your overall evaluation criteria. In some cases it could be a primary KPI, a secondary KPI and a guardrail KPI, but at minimum you need to have a primary KPI to say that again this is what will decide whether I'm winning or whether my experiment is winning or losing.

Ashley Stirrup: Yeah, makes a ton of sense and it's funny how it sounds so basic to like make sure you've got a clear hypothesis, a clear outcome specified. But you hear a lot of stories about people running the test and then looking for a way to show that it was a winner after the fact. Yeah.

Gaurav Sethi: And if I could just interject and add one more thing. So building an experimentation platform is just half of the job. I think if we put a good set of engineers together, maybe in six months you can build a decent enough experimentation platform. What takes a lot of effort, or at least from my side, took a lot of effort was educating teams on the best practices of experiment. And that is what took most of the effort or time from my end.

Ashley Stirrup: Yeah. Yeah, I mean that is a consistent theme across guests here on the show that you know it's i a lot of functions will talk about the importance of culture. I think it's particularly important here 'cause not only Do you want to people to think every time they build something, I should test that? But you want them to understand how to do it in a way that's got rigor. You want like one team that learned something to make sure all the other teams learned it too, so they don't have to keep running the same experiments or making the same mistakes or what have you. So it's you know, so there's so many different ways in which culture can have a huge impact on the impact of a experimentation program.

Gaurav Sethi: Yeah, I agree with that. Culture is very big when setting up experimentation.

Ashley Stirrup: Yeah. Were there any things in particular you did at PayPal to kind of address one of those opportunities?

Gaurav Sethi: Yes, I think one of the things that I was very consistent with was educational sessions, monthly, bi-weekly educational session. And I was pretty happy to see the attendance there as well. I think pretty much every other week I saw more than 100 people joining that call, trying to learn about it, and starting to implement some of those best practices as well. Then having regular office hours was very helpful. And I had a pretty good audience that would come in and join the office hours to be able to answer. And then I was also working on building an LMS program so that we can, as soon as a new product manager or somebody joins, they can just take like a basic experimentation one-on-one training so that they know at least the basics of running an experiment. And not just for product managers, but also for engineers on how you should be instrumenting, how you should be implementing as well. Just a little bit more technical as compared to product managers.

Ashley Stirrup: Yeah. Yeah, it makes a lot of sense. How about the kind of executive support and engagement in the experimentation program?

Gaurav Sethi: support from executives and one thing which was very interesting at PayPal was that we did not really, I had to prove the value of the platform every single year. It was not like it's a homegrown platform so everybody have to use it. So what I had to do was I had to make sure that every year we provide what is the value that it's bringing in. Number one, I did buy versus build analysis pretty much every year talking to a lot of the vendors. So that was an analysis or a report I was providing to our executives pretty much every year. we did see, like any other organization, there were most of the executives that wanted to have a data-driven decisioning and a very strict rigor around that. And I did get support from them as well. But there were some executives who were more worried about what their goals were and what they wanted to achieve. It was like a mixed bag, but all in all, I had good support from my leaders as well on building the platform.

Ashley Stirrup: Yeah. Yeah, it makes total sense. it's one of those things that if you look at experimentation with just your experimentation lens on, it's like, of course we should do this and every time and but you get into the larger business, you've got a variety of different priorities. I need to ship this new thing, what have you. but the opportunity to drive growth through experimentation is just so massive. So

Gaurav Sethi: I agree to that. And sometime I remember talking to Ronny Kohavi once and he introduced me to the term HiPPO effect. Sometime that HiPPO effect would show

Ashley Stirrup: Mm-hmm.

Gaurav Sethi: up in some of our conversations as well. But all in all, I noticed that most of our leaders were very much aligned with the idea of experimentation. The only concern they had was should we invest in a homegrown platform or should we go and buy an off the shelf platform.

Ashley Stirrup: Yeah, makes a lot of sense. So, you've already talked some about the technical challenges around running an A/B testing program, particularly at the scale of a product like PayPal's. maybe you could say a little bit more about some of the areas that you invested in.

Gaurav Sethi: I think one of the biggest area that I invested in was this 25 % dilution I gave between exposure and the assignment event. Because just to give you a little bit of context, like any other technical organization, a lot of our experiments or UI is driven by server. So what that means is the server is rendering whatever. So whatever is deciding whatever is to be rendered to a user on their client devices. So for engineers, it will. very difficult it was kind of almost next to impossible for them to go and change the base the most of their SDKs and code and everything so that when a server sends an insight sorry server sends a request to say Gaurav should see control or treatment the client has to log that as well whether Gaurav actually saw the control or treatment and there was this connective tissue that was missing and the innovative solution to that was an exposure event where when you're setting up an exposure you go and define an exposure event which is going to be unique to your program or unique to your project and you can always have use same as well. But you send you log that even as soon as you know that the user is going to be part of an experiment or as close as possible you are going to be part of an experiment. That actually was one of the biggest thing because I had to talk to like I had to align on many different teams data scientists because they already had built their dashboards and everything thing that was using the assignment data. The engineering teams because for engineering teams assignment was very easy the platform is doing the logging they don't have to worry about additional events or logging so if they had to do this they had to change the foundation of their features as well so that was a little bit alignment they are required and then finally the product managers as well because this would take a little bit extra time when they are launching a feature when we are on a very tight timeline or a tight schedule. For product manager it was difficult because this is not their primary thing they are looking for. They're looking for feature to be shipped faster. So

Ashley Stirrup: Yeah.

Gaurav Sethi: this assignment versus exposure was one of the biggest thing that I had to deal with. Other than that, I think the other thing that I worked on was building like a repository for an experiment. whenever the experiments were finished or run, teams across, let's say, check out multiple teams who are running multiple experiments, they consolidated all the reports for our leaders to be able to review them. That was a huge manual effort. So aligning everybody on how would we make a repository on the platform where you can just enter all the information, generate an automated report, and just share that with your leader, that saves easily 10 to 20 hours for two or three data scientists.

Ashley Stirrup: Yeah. Yeah. I can see how that'd be very powerful. And going back to your exposure piece of it, yeah, if people actually have to go make changes to the product to make sure you're tracking the right exposure events, that's that's a much bigger ask for people, right? Yeah.

Gaurav Sethi: where the exposure event method that we introduced was kind of a good solution because now they didn't have to make changes, too many on their side. At the same time, all they had to do was just log that one event and everything else will be done behind the scenes by the platforms. The data pipeline that we had built, it would run all the fuzzy logic and all the connections with the rest of the data set to be able to come up with the right analytics.

Ashley Stirrup: Yeah, yeah. That's that sounds like the right way to do it. in terms of kind of bringing in transactional data, you you mentioned the importance of that. were there were there a lot of challenges around how you did that?

Gaurav Sethi: Yes, that was again one of the obstacle to trying to figure out how do we generate ROI on the experimentation platform or a specific experiment. Because as you can imagine, since our instrumentation was broken, everybody was using different way to first of all log the data. And then the transactional system has its own way the way it's logging data and PayPal had so many people acquired almost, I think 10 brands between 2015 and So they had their own standard structures. So first of all, making sure that all of them are normalized to one single structure, number one. The second problem was the definition of a metric. There was no canonical metric definition that was available. I remember at one time when I searched for in our catalog, search for TP, which is total payment volume, I found 11 different variations of that. So which one is correct? Every team was using their own version of that. So aligning.

Ashley Stirrup: Yeah.

Gaurav Sethi: on those definitions, making sure that the transactional grain and the customer grain aligns with the grain from the behavioral data ecosystem. Merging all of them together, they should run on a tight schedule, mapping it to all the exposure events and all. That was another really complicated problem that we solved.

Ashley Stirrup: Yeah, yeah. No, you can just tell by how you're describing it that you really need a data expert to go do that. They might n then need to go get a bunch of people's buy-in, especially when you're thinking about do I need this data on a per transaction level, a per day level, what have you. and just understanding all the different data requirements for different experiments.

Gaurav Sethi: That's correct. And I had pretty good partnerships across the organization who were able to help me. And the reason for that was before taking over the experimentation team, I was actually leading PayPal's data ecosystem. So there I had done a lot of hands-on work with how the data resides, where the metrics are, how the grain of the data is. So that helped a lot as well in building this specific problem.

Ashley Stirrup: Yeah, yeah, makes total sense. so maybe to wrap up, how do you see experimentation evolving?

Gaurav Sethi: That's an interesting question and the way I see that is it's going to be a two dimensional approach. One is how do we use AI on existing or standard A/B testing methodologies? Things like building agents that can summarize an experiment readout for a user or a product manager so that they don't need to wait for data scientists to explain all of these things to them. Things like when you're writing a hypothesis, it can compare against your repository and say whether... An experiment like this has already been run or something similar has already been tested and how does that work? Recommending what should be your primary metric, what should be your target segment, all of these things, this is where AI would come in very, very, very close to helping users and guiding them on setting the experiment. On the other side, the second dimension, the way I see that is the experimentation for AI, which means that how are we now going to experiment or A-B test AI, be it LLM models, be it agents, be it any of the other AI features that we are going to build. And the reason for that is... Our traditional methods were pretty much deterministic. So you have a button. There are only three things that can happen. Either the button was rendered. If it was rendered, either the user clicked on it or they did not click on it. Whereas with agent, you don't have that determinism. An agent might finish the job, but it might end up burning maybe $100 worth of your tokens. An agent might finish the job, but it took, used the wrong tool to be able to access the data. An agent might not be consistent every single time when you're asking them, asking an agent to do the same thing. So things like this make it different on, make it different. So these methodologies, for experimenting with AI needs to evolve. And hopefully this will evolve as well as industry start catching up on not just building agent, now testing agents as well. Another interesting thing, or at least a thought in my head has been that I think by default, every agent that should go live in production should go behind a feature flag. that... Because if agent has access to a lot of data and it has access to a lot of context and it is not governed, that could actually be pretty, I would say, dangerous for some of the businesses. So in that sense, some of these things will evolve as industry make progress towards AI agents.

Ashley Stirrup: Yeah. Boy, you covered a lot there. Yeah, I couldn't agree more with your last point that, you know, as organizations are using AI coding tools to ship more features and ship faster, you also want to invest in things that help give you more safety. And feature flags is a great example of that, especially if you can tie it to performance metrics in your data warehouse so you can auto-roll back a feature that's hurting a customer experience. So the the more of those safety features you have, the faster you can go with rolling out new things. So totally agree on that. And then I think you hit on some really important nuances between the different ways AI can affect experimentation. Like I'm a big believer that there's this opportunity to take AI and automation and apply it to experimentation. So you know best practices on if you're doing this kind of test, you should have these kind of guardrails and this is You should set up an experiment. And AI can definitely help with that and apply it to different use cases. That that feels like a really good fit. Figuring out what your next idea should be, like that one I think we've it's gonna take a little longer before AI is a great fit for that. But then the last piece around A/B testing AI, you know, like we had Finn on there used to be Intercom that recently got bought by Salesforce, and you know, with a customer service chatbot. Did the person leave grumpy because you didn't really solve their problem, or did they leave because you actually resolved the problem? Like that requires you to take all the principles that you had for traditional A/B testing and kind of up-level your game, because you need to get better at defining what was the successful interaction and you know, how are you really gonna measure the difference of this prompt versus that prompt? but I think it's the only way to truly measure the benef you know, like is this faster model actually better than the smarter model or is this new model better than the old model? It's the only way to measure those things. Cause evals can only kinda do QA. They can't tell you the impact on a human user experience.

Gaurav Sethi: I agree with that. And I think that is something measuring the ROI for an AI agent or AI feature is going to be the next equal to web analytics. If you remember in 2010, 2012,

Ashley Stirrup: Yeah.

Gaurav Sethi: that analytics was big thing where we, everybody was trying to figure out how, how to measure the accuracy of any experiments that we running or any marketing campaigns we are running. think something similar needs to happen for AI agents as well, or AI features in

Ashley Stirrup: Yeah. Yeah, there's clearly some some that are very s straightforward, right? Like Zapier. They have an AI agent that helps you build a workflow. And did somebody actually build the workflow? Did they build it faster? Is it have they actually deployed it? Like those are easy to measure and therefore you can kinda test the outcomes. But if it's something fuzzier, like was that a good chat experience or not? Or did they write did AI write a good blog or not? Like those are harder and so they require additional thinking, new types of data, all that. So

Gaurav Sethi: Yeah, I think there will be like multi-dimensional approach we will have to take because it's not just one thing, yes or no kind of a thing like deterministic. So we'll have to look at it from multi-dimensions like, okay, did it accomplish the job or not? That will be the number one thing. Did it finish the job or not? Can it do this thing consistently? Can it produce the same results over and over again consistently? Did it understand user intent? Like if user asked for create an image for me, did it create the right image? where did it leave user with? And lastly, cost is going to become really critical as well. Let's say you say an agent that booked me a ticket. Yes, it books you a ticket, but it spent $3 worth of tokens just trying to book one ticket. So things like this, you will have to look at it multi-dimensional way. I'm really not sure yet how that would happen on an experimentation platform or experimenting with the agents. And I believe the industry will evolve and we'll have new methods to do that.

Ashley Stirrup: Yeah, I I couldn't agree more. and that's that's a great place to leave it today, Gaurav. You've been a terrific guest. I feel like we learned a lot, particularly around, you know, all the the things that happen behind the covers in order to build a great experimentation program. So thank you so much for coming on the show.

Gaurav Sethi: Thank you for having me, Ashley.

Gaurav Sethi is a Group Product Manager at PayPal, where he led Elmo, the company's homegrown experimentation platform, after previously running PayPal's data ecosystem. With more than eight years at PayPal and earlier experience at Hitachi Data Systems, he brings deep data and instrumentation expertise, a rigorous view of experiment measurement, and a track record of scaling a program to thousands of experiments a year.

LinkedIn
Gaurav Sethi
Role
Product
Industry
Financial Services

Subscribe to the podcast

Takeaways from this conversation

About half of PayPal's $180 million impact in 2025 was cost avoidance from features that tested badly and never shipped, which is why Gaurav argues there are no losing experiments.

S1 | E46

A six variant carousel test projected at 700 days was resolved by a multi armed bandit in 51 days, with a winner worth almost five basis points.

S1 | E46

Standardized instrumentation and canonical metric definitions are what let the analysis pipeline run without human cleanup, taking experiment readouts from four weeks to 24 hours.

S1 | E46

The exposure event fixed it. Fire an event as close to render as possible so the platform knows exactly when a user entered the experiment, instead of logging on page load.

S1 | E46

A win rate of 55% to 60% against an industry average of 11% to 14% was a tracking problem, not a performance one: assignment data carried 25% to 30% dilution and reached significance on users who never saw the test.

S1 | E46

Top takeaways from other favorite conversations

All Takeaways

Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric.

Go to S1 | E44
Theme
Growth
Role
Data Scientist
Industry
Business Tech
Featured
false

A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control.

Go to S1 | E34
Theme
A/B Testing
Role
Product
Industry
Consumer Tech
Featured
false

AI scales institutional knowledge, not just analysis speed — mining past experiment readouts to auto-generate new hypotheses turns your testing history into a compounding advantage.

Go to S1 | E12
Theme
AI-Native Dev
Role
Engineer
Industry
Marketplace
Featured
false

Most B2B product teams are feature factories. The fix is a top-down OKR system, and planning usually breaks in the connections between layers.

Go to S1 | E25
Theme
Culture
Role
Product
Industry
Business Tech
Featured
false

Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.

Go to S1 | E3
Theme
ROI
Role
Exec
Industry
Marketplace
Featured
true

Offline evaluation acts as a pre-filter for model velocity — Amazon's search team used golden data sets to cut hundreds of ML candidates down to 10 for live A/B testing, preventing wasted experiment slots.

Go to S1 | E12
Theme
Testing AI
Role
Engineer
Industry
Marketplace
Featured
false

Add real guardrails: track AI infrastructure costs, ethics/compliance, and inclusion metrics alongside growth KPIs.

Go to S1 | E6
Theme
Growth
Role
Exec
Industry
Identity / Gov Tech
Featured
false

Your customer and your user may not be the same person — building for HR specialists instead of the HRBPs who actually run talent reviews resulted in a feature nobody could use.

Go to S1 | E13
Theme
Growth
Role
Exec
Industry
Business Tech
Featured
false

Grubhub tests big bets by releasing to internal users or small cohorts first, so it can read business impact without disrupting core revenue and order metrics.

Go to S | E
Theme
Scale
Role
Product
Industry
Marketplace
Featured
false
The experimentation edge podcast logo with a picture of host Ashley Stirrup