Back to Podcast
A/B Testing
Culture
Scale

How Kargo turns losing experiments into competitive edges

S1 | E27
Jul 14, 2026

Summary

In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure.

Chapters

00:00 Welcome and introducing James Falzone
01:45 What Kargo does and how real time ad auctions work
04:45 Why experimentation is embedded in Kargo's culture
07:45 The three things every marketplace has to deliver
10:15 The experiment that failed: click optimization on third party demand
12:15 A bad result versus a bad experiment
13:45 Why different customer types need different signals
15:30 Putting "where did you fail?" on every retro
18:45 How experimentation evolves with AI
21:15 Better not bigger: the closing takeaway

Notable Quotes

"There's a difference between a bad result and a bad experiment. If you're not failing, are you really trying anything new?"

"The bad results didn't come from technical implementation. They came from a lack of contextual implementation."

"The same metrics and the same signals don't apply to every customer type."

"A buyer will never tell you what that value is. It's up to you through experimentation to figure out how they're defining value before they do it themselves."

"Better not bigger. Everything needs to be built on top of solid ML and infrastructure, but more people can come to the table with more ideas."

Transcript

The Experimentation Edge - James Falzone

===

James Falzone (GUEST): [00:00:00] there's a difference between a bad result and a bad experiment because to some extent, like if you're not getting those bad results, like if you're not failing are you really trying anything new, you know? So I think for us, like the biggest learning there was that- the same metrics and the same signals don't apply to every customer type.

Welcome to the Experimentation Edge, where product managers, data scientists, and engineers talk about how they make smarter decisions. I'm Ashley Stirrup, the chief marketing officer for GrowthBook, and in each episode, I'll sit down with an executive to unpack how they use experimentation and A/B testing to make better decisions.

This show is sponsored by GrowthBook, the open source experimentation platform leader. Now let's jump in and get started with our next guest

Ashley Stirrup (HOST): Hello, and welcome to today's episode. I'm excited to welcome James Falzone, Director of Product [00:01:00] Management at Kargo

James Falzone (GUEST): Hey, Ashley, thanks so much for having me

Ashley Stirrup (HOST): Yeah. Thanks so much for coming on the show. To kick things off, since Kargo is a little bit of an unusual business, could you describe what they do?

James Falzone (GUEST): Yeah, absolutely. Happy to. So Kargo is, and that's Kargo with a K. It's Kargo is an ad tech company and we are specifically an ad tech marketplace. So you know, what that means is basically that we connect our buyers to our sellers, and our buyers are brands, advertisers, companies that we're all familiar with who want to place ads on publishers that have ad slots or inventory that where you normally see ads.

So our job is to sell that inventory and enrich it in a way that brands can actually target the audiences who are on those publishers. So when I say publishers, editorial in the open web, like when you're reading an article, um, CTV, social, and even most recently, OpenAI for Kargo. You know, I think that one thing that's very fascinating that people [00:02:00] don't always know when, when they hear ad tech and they hear about these marketplaces is that, these ads are actually placed via auction, uh, and specifically what's called a real-time auction.

So, being in New York, I'll use this example. if you're on, your phone and you're reading about how the Knicks won the championship, uh, and you're scrolling and you're reading about the amazing tip-in and you see an ad pop up, before that ad, before that page actually loaded , and there's that split second when you click the link and the page loads, it's not just the content that's loading.

It's actually a series of auctions that's taking place where that publisher is reaching out to a company like Kargo and saying, "Hey, I have this ad- advertisement opportunity in this geo for potentially this user or what I know about them. Who's interested in buying?" And it's Kargo's job to then submit a bid on behalf of an advertiser to make sure that actual advertisement lands on that spot.

So the winner of that ad, the winner of that auction actually gets to put their ad there, and all that happens within a second or a thousand milliseconds.

Ashley Stirrup (HOST): [00:03:00] Sounds crazy complicated. You gotta be so fast and so many different things to optimize on

James Falzone (GUEST): Totally. And there's so many different steps and intermediaries and decisions and levers to pull. But it's kinda cool when you kinda like peek under the hood and it's "Oh, wait. That's how that got there."

Ashley Stirrup (HOST): Yeah. It sounds like you live in the land of data science

James Falzone (GUEST): Y- yes, that's definitely one of the most prominent teams at Kargo, and I think just teams in the industry as a whole. I think one of the cool things about just being in ad tech is that, of course, AI in the last, five years or so, three, three to five years, has really boomed and, the advent of LLMs has just changed the way the world is gonna work.

Machine learning in ad tech has been a core part of the industry for the last 15 to 20 years. So you know, it's actually, the sophistication of the models and the testing that goes behind it is really in a solid place, and honestly a cool place to also be

Ashley Stirrup (HOST): Yeah. I'm sure [00:04:00] there's no shortage of intellectual challenges for you

James Falzone (GUEST): Yeah, no, definitely not

Ashley Stirrup (HOST): Yeah. Could you describe your role there?

James Falzone (GUEST): Yeah, for sure. I work on the marketplace product team, so specifically that kind of like middle layer between connecting the advertisers or the buyers or the companies that we all know to the publishers. We as a whole, our fundamental job is to just monitor the health or the heartbeat of the marketplace and make sure that it scales.

And when I say we, I don't just mean the product team. We are specifically known as the auction and outcomes pod, and that is a group of pretty large group of product managers, data scientists, machine learning engineers, and software engineers. And of course, there's your data engineers and your analytics engineers that kind of layer into it as well.

But the core group is very ML infrastructure, data science-heavy.

Ashley Stirrup (HOST): Got it. And so what does experimentation look like there?

James Falzone (GUEST): So it's interesting because you know, experimentation, I know that there are many companies who have experimentation [00:05:00] teams. And it's interesting because, and I'm not just saying this on behalf of Kargo, but I think it's an industry thing because the industry requires it. Experimentation is more like embedded in the culture of how we do things.

And the reason for that is because, you know, when you think of just the sheer scale of the industry, it's, it's impressive. I can tell you that on any given day, so an ad request, which is an opportunity, you know, a publisher loading a page, you know, or something like that, and that can happen up to 10 billion times a day.

We can place hundreds of millions of impressions in any given day. We have a marketplace with tens of thousands of advertisers and publishers. So there's just so much data that... And it's constantly evolving, that it's a requirement to experiment. So I think something that we take, as a culture component of Kargo, which is to always make sure that we're trying something new

Ashley Stirrup (HOST): Yeah. And I'm sure you have a variety of algorithms [00:06:00] across your platform and that you're constantly looking for an edge and trying out different models and things like that. Is that right?

James Falzone (GUEST): Yeah. Yeah, that, that's exactly right. I think it depends on what the goal is, Because being a two-sided marketplace you are constantly trying to find the equilibrium between everybody that's involved. So what may benefit one side may not benefit the other side, and then the intermediary as well.

So we have different layers of optimization and models and feature engineering and that are honestly constantly pulling at each other too, which makes it hard to like truly optimize. But there's a balance, but there's definitely multiple layers of complexity and features and testing and modeling

Ashley Stirrup (HOST): Yeah. And so does each team have its kind of own data science person or group?

James Falzone (GUEST): I think it varies. I think We try to keep things lean for the most part. So I think you know, and this is interesting because it talks about-- kinda goes into the topic of [00:07:00] like, when you think about experimentation, it's like, is it technical or is it contextual? You know, because when we think about like the makeup of our teams, we're mostly split up by function rather than team.

So the data science team and the MLE team and, you know, even like the systems or software engineering team, like they're verticalized to some extent in terms of like what kind of function within that flow of the auction or the flow of the marketplace they're trying to serve. Yeah. It's an interesting, it's an interesting dynamic.

Yeah.

Ashley Stirrup (HOST): that kind of makes sense 'cause I would assume that if you've got individual teams that are focusing on different parts of the transaction and they might have competing goals that you really need to look at the experiments holistically across those teams, yeah?

James Falzone (GUEST): Yeah, totally. And I think that's ultimately when you think about like a marketplace, right? And this kinda eventually sets the guardrails for experimentation as a whole. A marketplace, whether you're Kargo or you're like an Uber or an Airbnb, there are [00:08:00] just like certain fundamentals that always are true, which is that, you don't own the inventory pricing dynamics are constantly shifting, but there's always three things that you have to accomplish, and that is to provide revenue to your supplier, profit to yourself, you gotta make money, and then value to your client, to your buyer.

And while revenue and profit are like quantitative figures, value is not. Value is constantly changing. So a customer or a buyer will never tell you what that value is. So it's up to you through experimentation to figure out how they are defining the value before they're doing it themselves

Ashley Stirrup (HOST): Yeah. That's super interesting, especially 'cause you'd imagine that, some of your clients are probably really sophisticated and others are probably not. And so you might have better insight into the true value you provided than they have

James Falzone (GUEST): Yes, definitely. Definitely. And then, sharing those insights is also a challenge, an exciting one in itself, because sometimes, there's so [00:09:00] much complexity behind the internet itself that it can only be measured and experimented on to a certain extent

Ashley Stirrup (HOST): Yeah. Yeah. Boy, that's... We could have a whole episode just on that topic 'cause I know as somebody in marketing that, I spend a lot of money on LinkedIn and Google Ads that it's really hard to track all the way down. Somebody doesn't see an ad and buy two seconds later, right? They do some research, they see three more of your ads, and maybe they talk to a salesperson, who knows, right?

James Falzone (GUEST): Maybe they go to ChatGPT now,

Ashley Stirrup (HOST): that's right. Very good point.

James Falzone (GUEST): Yeah. So it's all very, it's all evolving

Ashley Stirrup (HOST): Yeah. Could you tell us about a specific experiment where you had a lot of learnings?

James Falzone (GUEST): Yeah for sure. And I think actually going back to the culture of Kargo, of the focus on experimentation, I think we probably learn more when we fail. So I'm happy to share an example of an exercise, an experiment that we ran where we failed at first. And basically going back to the business model of Kargo.

So [00:10:00] we are the ones that place the ad on those publisher pages, right? Or when you're watching TV via streaming or something like that. And we have two types of demand. It's to some extent an oversimplification, but for the purpose of this, two types of demand. We have what's called direct demand, which is when the advertiser comes directly to us and says, "Hey, Kargo, we wanna work with you, help us do this."

Sure. But then we also have third-party demand, where actually the advertiser prefers to go to another party, and that party comes to us to tap into our inventory. So it's almost like B2B versus B2C, but always just B2B. Now what we had was our team basically launched an optimization strategy for clicks, right?

We wanted to make sure that we could optimize a client's budget to get as many clicks as possible 'cause that's the start of the journey. And we had a model. We were able to build that model with the appropriate features, predict within five milliseconds the likelihood of the click [00:11:00] occurring.

We were then able to, in real time, change the actual amount that we would bid given various parameters and whatnot, and we saw amazing improvements when we deployed that actual, what we call bidding strategy, with campaigns or, with campaigns that are direct demand, the customers that came straight to us were buying.

So we thought, "Hey, why don't we apply this same strategy to our third-party demand as well?" So the experiment that we started to run was pretty much we picked up the tech and the rationale and the reasoning, and we were like, "Let's put it right here." And we failed. It did not work. It was like, when I said earlier about the three things that you have to achieve, revenue to your supplier, profit to your...

it's just a bad results across the board. But I do think that, there's a difference between a bad result and a bad experiment because to some extent, like if you're not getting those bad results, like if you're not failing are you really trying anything new, you know? So I think for us, like the biggest [00:12:00] learning there was that- the same metrics and the same signals don't apply to every customer type.

And, it forced us to go back to the drawing board. You know, we had to, retouch our algorithm, retouch our models, review the features, review the outputs, review the normalizations of those outputs. And we eventually got there. You know, we eventually got there. But I think the learning was twofold.

It was The bad results didn't come from technical implementation, it came from a lack of contextual implementation. So I think it's just really important to make sure that your metrics and signals when you are testing are always business-driven. So I think that was the most important thing, and then the fact that, we operate in a multi-sided marketplace.

So that means that you could do everything right, but the market is gonna react to you as well. So there's always that to consider as well, which is basically okay, now we have external demand reacting to our changes. How is [00:13:00] that going to change what they do? So mul- multilayered. We were able to achieve the same kind of 25 to 40% increase in performance that we were really hoping for eventually.

But it was a good and valuable lesson that

Ashley Stirrup (HOST): Yeah. Boy that's a massive change. That's not a 1% or 2%

James Falzone (GUEST): Yeah. No, it was worthwhile

Ashley Stirrup (HOST): And were there things that you learned about the context from the third-party customers versus the first-party customers?

James Falzone (GUEST): Yes, actually. So that's a great question because I think first party you have access to a lot more data. You've a lot of visibility in terms of, I think, just simpler things like, what geographies they're interested in serving their ads, in what audiences, what kind of customers they wanna reach, the dates, how much they want to spend.

We have an idea of a budget. Whereas with our customers there's a lot of idiosyncrasies that, because if you think about it, it's not just them as customers, it's also their models that are now interacting with our models. So it's how do their models, what are the [00:14:00] idiosyncrasies of those?

And it turns out that different customer segments actually react to different types of data and different types of inventory. We eventually had to get to a point where we had to customize the model based on, the verticalization of those customers. But it forces you to dig deeper into, okay what do they really want?

' Cause they told us this, but their models are not reacting to that. So what are they really after? So it's fun. It's like

Ashley Stirrup (HOST): Yeah.

It's funny, I feel like you're giving us a little glimpse into the world when we all have our own AI agent doing our shopping for us and,

James Falzone (GUEST): And they're just there arguing with other agents. Maybe we should try to build an agent that can haggle online. Yeah.

Ashley Stirrup (HOST): Yeah. No, that's not really my best alternative. This is my best alternative. You need to lower your price

James Falzone (GUEST): Yeah, five cents less will do, yeah.

Ashley Stirrup (HOST): Yeah. That's so funny.

So in general, how do you try to extract as much learnings as possible from a losing experiment?

James Falzone (GUEST): That's interesting. I mean, I [00:15:00] think it's funny because what I might tell you, what I'm about to say might contradict itself to some extent. Because my honest opinion is that I really believe in tackling experiments as teams. Because, you know, the way that we run our pod is it doesn't matter really what team you're on.

You're a person first. You know, you're not an engineer, a product person, ML-- um, data scientist, whatever. So sometimes the best ideas can come from outside sources. Of course, everyone has their role. You know, as a product person, it's very important that you set the stage from a business perspective. But I think, you know, just being open about, the experimentation and the discussion points is, is key.

One of the ways we do that is when we have our retros. We actually have a little section that is like, "Where did you fail?" The sprint. It's actually kind of fun, and everyone talks about, "Oh, like, you know, I did this," or, "I did that." And so I think, just maintaining that dialogue and openness is [00:16:00] how you learn because I think analyzing the data and the table stakes and now that we have access to, you know, the various AIs that can help us like I think that's table stakes, but it's how you can then, pepper your creativity on top that changes the game.

But the contradiction though is that one of our goals is to really be an experimentation-driven company. You shouldn't need a team to experiment. You should be able to run an experiment as one person. And I very much believe in that mission. But do you lose the ability to talk to people?

So that's something that we're working with, which is like we're trying to work through is how do you allow that autonomy and that speed and that, flexibility and startup mentality while also maintaining the importance of creative thinking in a group and bouncing ideas off of each other.

Ashley Stirrup (HOST): So much we could unpack in that. I really liked your point about where did I fail? And I... Did you say you're doing that on a weekly or a monthly basis?

James Falzone (GUEST): Bi- biweekly.

Ashley Stirrup (HOST): And so yeah, 'cause basically you're making it okay to fail.[00:17:00]

James Falzone (GUEST): Yes

Ashley Stirrup (HOST): Like you said, where all the learnings come from is the losses, right?

You... if we never lost, we'd never need an experiment. We'd just ship everything we

James Falzone (GUEST): Yeah. Yeah

Ashley Stirrup (HOST): so the losses is where you go, "Okay, that... My assumptions were wrong or my guesses were wrong." And so like how do we learn from that? And so making it something that you explicitly talk about, that's creating a space for that shared learning to happen.

So that's pretty powerful.

James Falzone (GUEST): absolutely

Ashley Stirrup (HOST): And you're running a lot of experiments with a lot of variables too. Is that right?

James Falzone (GUEST): Yeah. Yeah. Yeah. A lot of variables. So much so that if we tried to record everything, it would cost us more than the benefit of the product that we tried to launch. So yeah there's a lot to it. Yeah.

Ashley Stirrup (HOST): Yeah. Yeah. Wow. That you must be doing some... You must have some special secret sauce behind all that in order to scale like that.

James Falzone (GUEST): [00:18:00] We're trying our best. We're trying our best.

Ashley Stirrup (HOST): Yeah, that's great. As a final question here how do you see experimentation evolving at Kargo?

James Falzone (GUEST): Well, I think that

We're lucky because we do have so many levers that we can pull. I think that the ad tech industry makes it hard to some extent to experiment because we have very strict latency limits, or latency constraints. You know, everything has to operate essentially within that one second, and there's so many steps.

But because we have so many levers, it means that there are so many possibilities. So I actually think that, you know, it's hard to think about how do things change in the future without, of course, mentioning AI. And, from my perspective, I think what AI is doing is that it is unlocking more opinions and more experimentation opportunities because people who otherwise would not have access to the code or access to the understanding of the code now do.

[00:19:00] Now if I go back to the constraints, LLMs cannot operate in the ad tech world just yet. The latencies. They just, they're not fast enough. So-- But they can operate at an orchestration layer. So back to what we were saying about having our agents negotiate prices for us, we could have an agent running tests for us, and that's what we're thinking going forward.

So I think, things are about to get hit a larger scale which is an ama-amazing thing. We can't treat it like a magic bullet because everything needs to be built on top of solid ML and infrastructure engineering. But I do think more people can come to the table with more ideas, and I think that's pretty exciting.

Ashley Stirrup (HOST): Yeah. Yeah, that's great. I think you're hinting at a lot of different things. Like with AI, it's interesting. We're all living through the world being totally disrupted,

and some of us are better and worse at dealing with that. But all of us, we're still learning, right? We don't know where is AI gonna be good [00:20:00] and where is it not.

And I think we've, we're seeing that like you said, with the technical, like suddenly people can get access to data, maybe it's generating SQL statements for them or something like that. And so it's lowering the barriers in some places. In other places, it's generating a whole lot of data that's not that useful or what have you, right?

Like AI can fall down.

And yet when you listen to the thought leaders at an Anthropic or an OpenAI, they challenge you to build for the next model so that when the next model comes along, maybe it's not very good at doing this thing today, but maybe tomorrow it will.

James Falzone (GUEST): Totally

Ashley Stirrup (HOST): A business like yours, it has to be right.

Like money is on the table. It could have a huge impact in, a few seconds, I would guess if it's going the wrong way.

James Falzone (GUEST): Absolutely. A mistake, within a minute can spend the entire budget of a campaign that you don't want it to. And I love what you're saying because it's true, because I think there's this push that we should all understand, which is, of course that is the [00:21:00] future and that's the direction that we're going.

But better not bigger, better not bigger, and I think

Ashley Stirrup (HOST): I love that. Yeah. And we're all going through this journey of is it good at this and is it good at that? And

James Falzone (GUEST): Yeah.

Ashley Stirrup (HOST): people you can get trying these things, experimenting, learning, getting comfortable the faster the innovation. So I think that's just a fabulous takeaway and a great note to end on.

Your business is very unique. I think about a lot of the other folks we've had on the show, and they're much more relatable businesses, like a Home Depot or a UPS or something like that. But what you're doing is really fascinating, and it was great to get a little window into that.

James Falzone (GUEST): Thank you. I really appreciate it. It was awesome to be here and thanks again for having me.

Ashley Stirrup (HOST): All right. Thank you so much

[00:22:00]

About James Falzone

James Falzone is Director of Product Management at Kargo, where he works on the marketplace product team connecting advertisers to publishers through real-time auctions that resolve in milliseconds across up to 10 billion daily ad requests. He champions embedding experimentation in company culture rather than a siloed team, and believes bad results usually signal missing context, not broken tech.

Role
Product
Industry
Business Tech

Subscribe to the podcast

Top takeaways from our favorite conversations

All Takeaways
No items found.
The experimentation edge podcast logo with a picture of host Ashley Stirrup