Back to Podcast
A/B Testing
Testing AI
Velocity

How Fin does 1,000,000 A/B Tests in 24 Hours

S1 | E28
Jul 21, 2026

Summary

On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.

Chapters

00:00 Welcome and what Fin actually does
02:00 How Fin became Anthropic's first line of support
02:30 Why Fin sells resolutions not deflections
06:00 Owning the stack with custom models
10:40 Pedro's path from fuzzy logic to AI
12:55 Why A/B testing is the only gold standard for AI
15:50 Do no harm testing on every change
18:00 The latency experiment that shocked the team
27:30 When more context made Fin hallucinate
30:15 Win rates and the future of AI driven experimentation

Notable Quotes

"Unit tests, like traditional software unit tests don't work with AI."

"Whenever we have any kind of like question dilemma, we just put it to the test. We just run an A/B test."

"This is so counterintuitive that my manager forced me to run a confirmatory experiment, check the data, make sure that there's no sort of a selection bias."

"AI development is inherently very uncertain very experimental, because you never know how much ground you have covered."

"Decision-making is always gonna remain fundamental, and decision-making can only essentially be done through experiments."

Transcript

The Experimentation Edge - Pedro Tabacof ===

[00:00:00]

Ashley Stirrup: Welcome to today's episode. I'm excited to have Pedro Tabacof, Principal Machine Learning Scientist at Fin, which was formerly known as Intercom. Welcome to the show, Pedro.

Pedro Tabacof: Thank you, Ashley. I'm glad to be here.

Ashley Stirrup: I think today's episode is gonna be particularly exciting because Fin has been really at the forefront of applying AI to customer service and a lot of other great use cases. Maybe to kick things off, Pedro, you could just tell us a little bit about the kinds of things Fin's been doing in that area.

Pedro Tabacof: Sure. I think a great way to explain it by just taking a historical perspective. Intercom started like almost 15 years ago as a customer service software. It has evolved a lot throughout the years, but a massive inflection point has been the release of Fin three years ago or so.

It was so important that we changed name from Intercom to Fin, which shows how Fin has become really the fundamental product of previously Intercom. Fin was [00:01:00] released aft-right after the ChatGPT moment when LLMs went from being a more, let's say, academic interest to a technology that can actually solve real-world problems, provide like real business value.

Not to say that LLMs were useless before, they were definitely useful. Intercom was using LLMs for some applications, but ChatGPT just unleashed like this w-new wave of A-AI to the world. And Fin was one of the first products really in the world to capitalize on this new wave. We launched early twenty-twenty-three using GPT-4 while it was still like in a public beta.

One of the first companies to use the GPT-4 API back then. Fin as a product has evolved a lot. We have always been AI agent for customer support but at first was mostly geared towards, I would say, more simple, more basic informational questions. Now, Fin can take actions, can follow hundred-step procedures Fin can follow guidance can use human help has memories, can read images, can send [00:02:00] images, and so on.

So we went from being almost like a Q&A to a full-fledged agent that can do almost anything that a human can do as well. So we're really happy with how Fin has evolved. It became the most important product for Intercom. It changed the company's trajectory and name. It's one of the most valuable AI products out there.

We have crossed a hundred million in ARR. We an interesting fact that we are Anthropic's first line of customer support. We have over ten thousand customers, many famous AI companies are Fin's customers. So yeah, we're really happy with the progress and relevant to our conversation is that Fin is what it is today from a lot of experimentation, from hundreds, thousands of A/B tests that we have run throughout those three years.

Ashley Stirrup: Yeah that's a pretty amazing story. Another really interesting point is that your business is based on you get paid per, basically deflected trouble ticket type of thing, right? Like the number of chats that you actually are able to resolve without [00:03:00] a human.

Pedro Tabacof: Yeah, we call them resolutions, not deflections, because deflections has a stigma which you can just know fob the user off. It's kinda common, unfortunately, in customer support, either human or AI customer support. A lot of deflections happens, but not real resolutions. Ultimately, what we really care about internally are real resolutions, actually resolving the user's issue.

And unfortunately, you cannot measure real resolution exactly, right? Because how can you truly know if a user has been helped, for example, when they go away without providing any sort of feedback? But we still try not just to measure deflections, but we use other proxies. For example, if they left after negative feedback, we don't count as a resolution because clearly we haven't helped them.

If you only engage with the user without providing information, without providing anything to them, was just like a hi, bye situation, we also don't charge because that's not a real resolution either. So there are many ways that like a deflection is not a resolution, and ultimately we care about resolutions [00:04:00] that what we track measure, and as I said we sell them.

We are a resolution-based business model, outcome-based. That was one of the first AI businesses to do so actually. Back then, three years ago, this was massive unpopular or it is very like contentious. A lot of people thought that you couldn't just work a product by offering selling outcomes.

You had to sell usage. You had to sell like a commitment, spend, or seats but not outcomes, because outcomes can be quite variable. Different customers have different resolution rates. This can fluctuate over time. By improving the product, you can increase resolution rate and so on. So the expectation was that the market was not ready for this, and maybe it was in twenty three, but our growth shows that it has been really like a wise choice

Ashley Stirrup: Yeah. The thing that's so powerful about it too is that the better you get, the better it is for your customers and the more money you make. So you're just totally aligned on investing in the things that help your customers.

Pedro Tabacof: Exactly. Whenever [00:05:00] we launch or run an AB test, for example, we are gonna look at resolution rate as one of the primary metrics, and we know that if we can increase resolution rates, we are, it's a triple win really. The end user wins by getting help in their customer support query or tickets so we have a happy user.

Our customer wins, the business that buys win-wins by just having their own users being happy and not also using their human support, so they can save costs and focus on the hardest problems, hardest queries. And we win by charging more resolutions. Of course the trick, the problem here or the danger is Goodhart's law, right?

That whenever a measure becomes a target, it ceases to be a good metric, or so it goes. And the risk here is that what if we for example, only increase the deflection side of resolutions but not the real resolutions? That's a danger. We have to be aware of this. That's why we have a lot of, let's say, guardrail metrics and complementary metrics to ensure that an increase in resolution is not increase in just [00:06:00] deflections, but also an increase in real resolutions.

For example, by measuring positive feedback, customer satisfaction, so on.

Ashley Stirrup: Yeah. Yeah, that's a really great point. I love it that it's not even just a double win, it's a triple win. And you guys have also been not just at the forefront of applying LLMs to businesses, but actually developing LLMs themselves. Is that right?

Pedro Tabacof: Correct. This has been ongoing for, I would say, at least a year. We took baby steps. We started, let's say, with the easier problems at first, though not easy in absolute terms, but easier than competing with let's say Anthropic or OpenAI. Initially, we invested heavily in our RAG Retrieval Augmented Generation part of the pipeline.

So we deployed our own embedding models, our own re-rankers that were trained with our customer support data to maximize resolution rate, retrieval metrics, and so on. And then we started to slowly transform other parts of Fin into custom models because then you [00:07:00] have stronger-- more control of what's happening.

You can train exactly what you want. You can choose the targets or the labels that better align with the business. You can use reinforcement learning to even use a more indirect reward to attach the model what it wants. You can control like the trade-off between like cost and latency, for example, and volume.

And ultimately, it's a strategy where you own your stack like vertically end-to-ends. We don't own the GPUs, unfortunately, but everything else is under our control which makes us very resilient to whatever happens out there we can run things on our own. Fin is a very complicated beast.

It's not like a one single prompt, one single LLM call. Fin is a collection, like a very complicated agent with dozens of LLM calls for different purposes. Some are very simple, maybe some classification, routing tasks. Some can be extremely complicated, like planning or answer generation one. And Fin is a mix of custom models and third-party models, and even our customers might have different models [00:08:00] depending on our contract with them.

But we really have doubled down on this approach or this thesis that Fin should be mostly powered by custom models as much as we can to really own the stack. And this has actually paid off handsomely because we managed to like get, I would say, broadly speaking, better results with, faster and cheaper models.

In other words, by leveraging like smaller models than probably what third-party LLM providers provide to you we managed to get like a even better quality that came from like loads of A/B tests, so like we are very confident in the results with smaller models, which means you can run them more cheaply or faster or both.

Ashley Stirrup: Yeah. And then the other interesting dimension for our listeners to keep in mind is it's not like you're an e-commerce store and you're just trying to optimize how people find red shirts or something, but you're needing to build a system that then works for every one of your customers, and each of their data [00:09:00] sets and businesses and customer service experiences are different.

And so you're really building a platform that then gets applied to lots and lots of customers.

Pedro Tabacof: Yeah, we have I don't know the exact number, but in the order of 10,000 different customers across all sorts of different businesses from pharmaceutical, medical, financial to all sorts of different e-commerce, to a lot of B2B software. So Fin has to be able to handle all of those different types of problems.

So a lot that we have invested throughout those three years has been on customization, so letting our customers configure Fin prompt Fin however they want. And the key here is that you don't wanna outsource all the prompt engineering problems to the customer, because our customers might not be technical enough to know, and even seasoned AI engineers might struggle with prompting or telling LLM what to do.

And what if the LLM changes, right? Because we're constantly updating models, whether our custom models, whether third-party models, all of them have a lifetime, a shelf life, so [00:10:00] we have to be constantly evolving. And how can you make sure that whatever the customer configured or explained to Fin is gonna be working with like new models, new architectures?

And the way is really twofold. One is by creating a system that's really generic almost like a meta system where instead of trying to like, for example, optimize refunds in particular, like how to provide refunds, you just optimize for policy or process following, and then refunds just one instance of a policy or a process.

And if you can ensure that your process on average are followed very well and you track this with like offline metrics, online metrics and A/B tests, then you can guarantee at least statistically speaking, that your customer refund process is gonna remain valid for the foreseeable future for its lifetime.

Ashley Stirrup: Yeah. Yeah. Makes total sense. So we've established just what a rich environment you've been in and the expertise. I'd love to hear, have you tell a little bit about your own personal background and [00:11:00] how you, you came to Intercom and the fact that you started more in data science and have now moved into experimentation.

Pedro Tabacof: Yeah. Actually, my first job out of college, which was more than 12 years ago, was in AI, but it was old school AI where fuzzy logic systems. No one really talks about them anymore. You have to be probably over 35 years old or 40 to even remember fuzzy logic. It was very popular in the '80s, maybe '90s.

And it was really cool, by the way. We had a lot of interesting applications to optimize process control. But at the same time that I was doing this kind of work deep learning started to take off very slowly with ImageNet being won by a convolution neural network, the work of Geoff Hinton and Alex Krizhevsky and so on.

I got very interested in that at the time. And I started a master's work, some research work in deep learning. It's funny that deep learning it was, like more than 10 years ago I felt was actually too late to join th-this wave of AI because all the competitions had already been won.

Like, all the image competitions had [00:12:00] already been taken by deep learning. And what else is there to be done, it's funny in retrospect how wrong I was, but also how early I was in this journey. So I did some work in deep learning at the time on adversarial images variational autoencoders which were both popular at the time, maybe not so much nowadays.

But then I became a, I would say, a more traditional data scientist applying machine learning to problems such as credit fraud risk. I also work a lot in performance marketing, like lifetime value marketing attribution. And then after let's say this long detour, I finally came back to my AI roots at Intercom, now Fin working with LLMs.

At first, I was doing this more I would say a prompt engineer or AI architect role, and now this has evolved, as I explained, also working for our own custom models. So I've been training my custom models. I've done supervised training runs, reinforcement learning, and so on.

Ashley Stirrup: Terrific. And can you tell us a little bit about how experimentation works at Fin?

Pedro Tabacof: Sure. Ever [00:13:00] since I joined, like three years ago plus experimentation has really been like the cornerstone of how we really operate. Because it's really hard to otherwise know if your agent's really working properly or not. That is you can only like know if an agent works well to the degree you care about in a very broad statistical way.

Because we know that AI is not deterministic, right? Unit tests, like traditional software unit tests don't work with AI. Even if you try to like come up with all sorts of use cases or paradigms, your end user is always gonna surprise you, and there's always gonna be something new, like a change in the product or the world that's gonna make users ask different questions or behave differently So the only way to really know if your system, your AI system is working properly or not is at scale and statistically speaking, and A/B tests is just a experimentation, a perfect way to do we also do, of course, like a offline evaluation, and that can be an important input in the development process. But the gold standard for decision-making is really A/B testing. That's the only way to first measure [00:14:00] what you care about, right? With an A/B test, you can actually measure the business metrics, like resolution rates.

Second, you can get like a scale. Offline evaluation is usually done maybe one thousand up to ten thousand examples, but our A/B test now can easily get millions of samples in a few days. So our scale allows us to detect much, much smaller effect sizes. So our power, in other words, our power is massive with A/B tests, much more than almost any offline eval.

So because of those reasons A/B testing has really been fundamental. And over time, as we gain scale, I think we even double down more. In the beginning, when the scale was smaller the A/B tests had less statistical power. So they were, I would say, like less useful, broadly speaking.

But as we increased scale, they became much more useful because in literally one day, we can already be pretty confident ignoring seasonality, which is a big problem for some A/B tests. But one day has enough data to detect very small effect sizes. So the way that we make decisions here is through A/B testing.

Whenever we have any kind of like question dilemma, [00:15:00] we just put it to the test. We just run an A/B test. It can be a simple two-armed one, but we can run like multiple arms. We can run for one day, for one week or longer, and really depends on what you're doing. There are different reasons to run an A/B test.

Maybe you're exploring some hypothesis, or maybe you wanna make a actual product change, and you need to reach like a final conclusion based on business metrics. So we run the A/B tests for different reasons, but experimentation is like our bread and butter. So much like we literally recently hired an AI analyst for our team.

And this analyst is pretty much only dedicated to improving our A/B testing framework and stack, like our dashboards, our metrics our experimentation methodology and protocols, and so on. So we have now a full-time person just doing this meta work so that the scientists and the engineers can be more effective in their jobs.

Ashley Stirrup: And you're running a lot of experiments, right?

Pedro Tabacof: I would say we are always constantly running something between one to two dozen experiments concurrently. Right now I own like a couple experiments and my whole team probably owns like 15 experiments, and there's [00:16:00] maybe five, 10 more from other teams as well ongoing. We have run thousands of experiments by now maybe more if you count like some short-lived ones.

And we have probably shipped like hundreds of them to production. Hundreds of them were deployed in favor of treatments.

Ashley Stirrup: Yeah. And you even, basically you're testing any change that you make, right? That you're doing do-no-harm testing even in ti- in places where you're just making small changes, right?

Pedro Tabacof: Correct. Even some bug fixes are A/B tested. Of course, depends on the level of criticality, but i-if it's not business critical we are always gonna A/B test right a- A/B test instead of deploying right away. Sometimes very minor changes fixing a small prompt, let's say, like a comma or period, is gonna be A/B tested.

What I love today about the AI agent tools is that the bar to launch an A/B test, especially because we have a well-oiled framework, it's very easy for us to set up an A/B test. I can just tell Claude "Just launch an A/B test to evaluate X or Y," and I can look at the code, ma- see if it makes sense, and then [00:17:00] this makes us just churn through a lot more A/B tests than was previously possible.

And so we can be confident that the, our changes are always directionally correct. We can also look for interactions and make sure that nothing's degrading. Of course, maybe the risks now are not getting overwhelmed by too much information, too many p-values, and so on. So that can be tricky and requires sometimes human judgment to understand what really matters.

If you run dozens of A/B tests and each one has a dozen of metrics, a lot of them are gonna be positive. Let's say a p-value is gonna be below your threshold, whatever it might be by sheer chance, right? There's always gonna be, like, false positive findings, so we have to be careful about those false positives.

Really depends on the test, on the metric in particular, but sometimes we even run confirmatory experiments or we let the test run for longer not to look for a p hack but once we have established finding, sometimes we run for longer even after it's already we have reached the timeline the sample size, but we just wanna be absolutely sure it's a real finding [00:18:00] and not a fluke.

Ashley Stirrup: Yeah. Yeah, it makes total sense. When you're running as many experiments as you are, you're gonna have, false positives, and so it's important to be double-checking things. So you had a great example of an experiment where you had some good learnings when you were testing latency and the impact that had on the customer experience.

Pedro Tabacof: Yeah, we even wrote a blog post about it. You can read in our research blog. And what was surprising here first, this was a very unusual experiment because this was a very rare case where we're not really testing improvements. We are testing degradation in some sense by increasing the AI agent latency of its responses just artificially.

And the reason why we run this experiment is because at the time, leadership was really concerned about latency. They wanted to de-decrease latency and but we didn't really know what was the impact that they had on the business on our customers and users. We had this broad feeling that latency is bad, we should try to decrease it.

If you look over the literature of like [00:19:00] AB testing for latency like for e-commerce like, and for a search like Google, you're gonna find a lot of like papers saying even like 100 milliseconds like clear business impacts towards increasing profits. So we just decided to put this to the test.

Ideally, what we really wanted to test was decreasing latency, not increasing, but decreasing latency is hard, right? Because if we could decrease latency, we would have done so already. We wanted a faster fan. So the only way to evaluate this question was by increasing latency. So we had a clever experimental design where we increased latency roughly proportional to the natural latency distribution so that no one would really feel something's particularly off.

You cannot make the whole agent like much, much slower to the point like everyone notices what's going on because then that even contaminates the experiment itself. So we just mimic the natural distribution of latency. Of course, the average latency increased, but in a way that was not super noticeable.

And we just run this test for much, much [00:20:00] longer to collect enough data on small the small arms, the small fraction of arms of high latency. And what we found, which was completely unexpected almost shocking to like many people here was higher latency only saw good stuff happening. Of course, we expected things like deflections to increase with higher latency.

If it take too long, users are gonna bounce off. That's kinda expected, right? But what we didn't expect was that positive feedback would increase as well. This is so counterintuitive that my manager forced me to run a confirmatory experiment, check the data, make sure that there's no sort of a selection bias in the, test assignments or some kind of metric miscalculation And after a while, I was really confident that the result was a true result.

Of course, the effect size was not that massive, so you can claim that for all intents and purposes, it's not like a massive change in a way. But like we couldn't see any metric essentially going in the negative direction of worsening the experience. So we thought a lot about [00:21:00] this, and we even found some blog posts about this and here's our best guess.

There's some psychological effect at play here, where for an AI agent, which is very different from like e-commerce or search users expect some work to be done. Humans take their time. If you're engaging with a human, they're never gonna answer instantly. If a human answer instantly, you're gonna think it's a canned response.

They're just using some kind of macro. They really are not answering a question like from first principles or thinking about it. And a similar idea might apply to AI agents, where if it takes longer, it looks like it's doing more work, they... doing more effort. It just looks more human-like. And this makes users more prone to leave positive feedback and they answer surveys with

more happiness. And of course, some are gonna be deflected as well. So overall, there's only improvements. And very important before I wrap up, a caveat here is that we did not increase latency. Despite those findings, we still decreased the latency because our leadership still believed that for selling Fin it was very [00:22:00] important.

Imagine that you are doing like a demo to a key stakeholder, key customer. If it's too slow to answer, it just looks bad in a demo, any kind of evaluation setting. So we did over time decrease latency a lot. Maybe it's twice as fast as it was at the time. But now we know something important.

Whenever we decrease latency, we have this confounder effect that we might see things worsening just because we are reducing latency. So for many experiments that we targeted latency reduction, we created an artificial third arm to keep latency constant as to remove this confounding factor so we can focus just on the effect that we are looking for and not the latency impacts.

Ashley Stirrup: Got it. So what you're doing is you roll out a new feature and maybe it's smarter and faster. You try to test the smarter and the faster separately a bit.

Pedro Tabacof: Correct. Yeah, correct. That's the only way you can disentangle the different effects, yeah.

Ashley Stirrup: Super interesting. In general, how do you try to help people at Fin learn [00:23:00] from losing experiments?

Kind of what's your approach and how do you try to extract as much learning as possible from every losing experiment?

Pedro Tabacof: I'll say there are two ways. One is to provide a lot of I'll say debugging metrics. We have a set of metrics that we call the summary metrics, which we care about the most. But also we have over time built a collection of a long list of maybe thirty pages long of different metrics that can help debug and understand what's happening because they can be confounding factors.

Let me just give an example. The rate at which Fin asks questions to the users let's say clarifying questions, interrogating questions, this has a huge impact on some metrics. So if an experiment loses, let's say it worsens resolution rates a little bit. Is it because it's actually worse at real resolutions or maybe because it's asking fewer questions or more questions depending on how it goes?

It's important to know the root cause, and we have collected a lot of such things, which might not be of direct interest, but might be, like, mediate an important effect. So we try to really dig deep. And now another [00:24:00] way to do things which I really like honestly is just using AI agents.

Like we are heavy users of Claude here. So Claude can go... With experimental data, Claude can generate hypothesis and use the data to test a lot of hypothesis. You can have Claude create let's say, offline evaluation system of other LLMs use LLMs to judge to figure some-something's fundamentally different.

And then you can have Claude give you some kind of inputs for the next ex-experiments. And now we have this interesting cycle where you run experiments, Claude can evaluate the results for you and really try to come up with different angles and hypothesis. And whatever it comes up with, you can experiment again.

So you're not just relying on Claude's claims or hypothesis. We do not be fooled by the AI, right? Which can happen. We wanna make sure that the hypothesis or the claims are backed by another experiments, and we can keep doing this until we either give up or learn something and finally deploy the treatment of of the change.

Of course, gotta be careful not to, again, p-hack your way out of this. By [00:25:00] running a lot of experiments one of them is gonna end up positive. So we should be really skeptical of p-values in particular in those situations where you really try hard to deploy something

Ashley Stirrup: Got it. And so I just wanna make sure I kinda heard you correctly there. So basically what you're doing is you're using AI to analyze experiments at scale, come up with potential conclusions, and then retesting those individual conclusions. Is that the right way to say it?

Pedro Tabacof: For example, if we notice that the problem is because of increased Fin is asking more questions, maybe he shouldn't ask that many questions. You can have Claude, for example, research this problem come up with a more specific hypothesis or changes. Maybe change, for example, a prompt to make Fin ask fewer questions, evaluate this o-offline so that you can ensure that yeah, on those new, the new variants, the question asking is the same.

So we remove that confounding factor, and now we can test again and check the overall results. So Claude is really great at analyzing, poring through data, manually or automated, but what is [00:26:00] nice about Claude or other AI agents is that you can have it go through the data directly read through like answers at scale, because humans really don't have the patience or even the mental capacity to read through like hundreds of messages.

It's very hard to do so unless you really dedicate a lot of hours for it. But Claude can read through loads of answers and messages and examples, and then look for systematic patterns that then can uncover like a root cause or problem that then you can experiment on

Ashley Stirrup: Got it. So you're actually looking at the questions and answers getting, actually b- that are occurring on your customer sites and trying to understand the patterns there. Is that correct?

Pedro Tabacof: Yeah it really depends on what you're trying to measure or experiment with. But I'll, I would say that you can feed Claude pretty much anything really. But sometimes it does get raw data, sometimes it gets sanitized data depending on the situation. And the idea is that by having Claude read through the data even in raw form, it can find some patterns, some problems, and it can come up with [00:27:00] hypothesis.

We even have a brainstorming skill that can allow us to generate a lot of hypothesis. And then once we have a hypothesis, we can test this offline with some kind of like back test or LLM-based evaluation. And then we can go to an online experiment to confirm the hypothesis for real

Ashley Stirrup: Yeah. Yeah, 'cause I think there's a really im-important point here that's getting a little hidden, which is that maybe you've rolled out a new feature and that feature might be a great idea, but there might be an element within that feature that isn't, ideal for the customer experience. And so that element might be bringing the performance down to that.

And so if you can figure out which part of that new feature is actually hurting the customer experience and fix that, then suddenly you've turned a loser into a winner.

Pedro Tabacof: Yeah I have a good example. This was actually pre-Claude relied on complete human judgment and hard work to get this done. But it's an interesting one where we were adding more context to the answer, answer module. So we were providing more conversation history [00:28:00] context when Fin was providing answers.

And this seems like a very obvious thing to do, right? The more context you provide, the better, the more it understands the problem the user has. It was a very natural thing to evolve towards. And what we noticed through both offline and online experiments was that by adding more conversation history Fin was actually hallucinating more.

We have a very important guardrail metric on hallucinations that we measure with LLM judges, both offline and online. And we noticed that adding more context worsened hallucinations. But also, we saw massive improvement in positive feedback. So we had this mixed result. Way more positive feedback, which is great.

Users seem to be helped more, but a lot of them seem to be caused by hallucinations. For example, Fin could say things such as, "Yeah, of course we're gonna process your refund." This can be a very serious a fake promise. We never want to let Fin do it for our customers. It wasn't super common, but was still common enough to be, like, a [00:29:00] measure in our metric.

So this forced us to go back to the drawing board after the first experiment failed because the hallucination increase was just not acceptable despite the increasing positive feedback. So we tried... Iterated over the prompt multiple times. We tried to figure out offline all of the new patterns of hallucinations that were, like, being created by having more context more conversation history.

And then we addressed those in the prompt, and then we ran a secondary experiment saying, "Okay, what is the real effect of adding this context with those new instructions without the added hallucinations?" And yeah, the positive feedback reduced the impact, but it was still there. So maybe it halved the impact.

But we still had a sizable impact of increased positive feedback and now without any hallucination increase. So this has allowed us to move forward and deploy the new treatments.

Ashley Stirrup: Got it. And so did you keep the additional context then, and it was just about tuning the prompts to make sure it was using that context correctly?

Pedro Tabacof: [00:30:00] Exactly. I would say for most changes which fail, like when you're handling AI-based systems, LLM-based systems, almost always with the right prompting you can overcome like most of the challenges, not all. And in this case, it was exactly that, where by adjusting the prompt, we were able to like prevent the new hallucinations from happening, but still benefit from the added context.

So we went forward with the change. So Fin, like for a long time now, has the a much large, a much, a stronger, let's say, conversation history context without any kind of like a worsening or increasing hallucinations

Ashley Stirrup: Got it. What would you say your overall win rate is just approximately?

Pedro Tabacof: We don't really track this. Maybe we should. Just my intuition would be roughly around 20% maybe 30%. I think it really depends on the-- even the definition, like what is a valid test? What is the denominator here? The numerator, I would say it's easy, like the number of A/B tests that we have shipped treatments, that's somewhat easy to compute.

But the denominator is actually much [00:31:00] v-very tricky because, for example, what if you launch an A/B test and you realize something's wrong, like on the hour zero of the experiment or hour one? We have a real-time dashboard for like-- which we track like a error rate, some superficial metrics.

What if we find a high error rate and we stop the A/B test very quickly and relaunch with a quick fix? Does that count as one or two experiments? Or what if we have an A/B test which evergreen so we can compare like two hypotheses forever? There's some situations where this might make sense.

Should that count as an A/B test but we don't even intend to deploy it? We are just using the A/B test framework to measure, how the system's per-performing against some kind of baseline or fallback system and so on. So just finding like the right number of experiments to consider is difficult, but I would guess 20%, 30% would be a rough approximation.

Ashley Stirrup: Yeah. And yeah, whether it's 10 or 20 or 30, I think, for a lot of people who aren't experts in A/B testing, I think they implicitly assume that [00:32:00] 90 to 100% of their new features are winners. And so when they find out their win rate's below 50%, the-- for the people that, have an open mind, they re- it's a humbling experience and it causes them to say, "Okay, I need to test more.

I need to learn more," all that. How would you say, How would you describe the mentality at Fin around, is there pressure to get winners or are people pretty open to having a lot of losers 'cause they know that, you have to do that in order to learn?

Pedro Tabacof: I think the culture here is very favorable towards experimentation. The leadership understands that AI can be very tricky to get right, and they don't want just us to make vacuous claims of, "Oh, yeah, Fin has a new feature." For example we added maybe last year or so, the ability for Fin to read the images and also the ability of Fin to send images to users.

Those are two different features that were both A/B tested separately. And it's not something you can just do overnight. You could, right? You could just get a multimodal LLM, put in production, and just Start to reading images and writing lou- [00:33:00] writing down some transcriptions and then choosing some image that seems plausible.

You could just ship it literally one day with all the LLMs that we have available, the APIs, coding agents, and so on. But that's not what we want really. What we want is actually a high quality multimodality feature. We want, for example, whenever Fin is reading an image not to hallucinate content of the image, for example, not to make up a number, like error code that might be in the image.

And this might not come from... What if the image is blurry, right? Should Fin try to infer the number or should it be more conservative? Those things are hard to really think at first. You have to iterate a lot of times to go through a lot of corner cases, different situations, and really dial down what you care about.

That means that AI development is inherently very uncertain very experimental, because you never know how much ground you have covered. You never know if the next prompt tweak or the next LLM you try is gonna be the one that's gonna crack the problem. Even the example I provided previously [00:34:00] on the hallucination increase with more context.

If we hadn't measured this properly in the first experiments, we would have deployed a terrible feature that would have caused an increase in hallucinations. And when we saw the failure, we also didn't know if we could find a solution. We didn't know for example, if the next iteration would fix hallucinations.

We didn't know whether the positive feedback increase would remain as a part of it. All of this is empirical, and our leadership understands this, and they would much rather have high quality product with fewer hallucinations, higher real resolutions, higher real positive feedback than just a crappy product which just claims a lot of stuff has a lot of long feature list but then works in a not so reliable way.

So they don't really know I would say like much about our day-to-day experimentation. They don't know much about the A/B tests we have ongoing, right? They just take a higher level view, but their higher level views are like, yeah, it is uncertain, requires a lot of iterations, a lot of hard work, just the grind.

But in our end, we do the [00:35:00] grind through experimentation. And luckily, my manager, who's now the chief AI officer, he comes from an experimentation background. He even sold a company he founded many years ago to Optimizely. He worked at Optimizely for a while, so he had a strong bias towards experimentation, and this actually made our life much easier.

He understands all of the reality behind A/B testing, and so we always had also this protection from our AI group lead, from day one.

Ashley Stirrup: Yeah. I think that example you just gave of reading images is such a powerful example. We all know that there's been tons of "Oh, the prototype, it looks so awesome," right? Until you actually care about what's being written and that it needs to be read correctly every single time. And it just makes so much sense that with something like that, where you can have so many different kind of corner cases, a blurry image or what have you that the only way you're going to get to an excellent product is that you're measuring and testing and you're looking at multiple different [00:36:00] dimensions of the product experience and iterating on those.

So it's a real testament to your commitment to experimentation and how that's led to such a better product. As we wrap up, how do you see experimentation evolving at Fin?

Pedro Tabacof: As I said earlier, we have now a dedicated AI analyst just work mostly on experimentation so that we are even better, let's say, well-equipped in terms of methodology, metrics data vis, dashboards, and so on. He has already done like tremendous work making my life as a consumer of A/B test, not just mine, but my team's my life's much better now.

What I do almost every day after I wake up is check this Slack channel where we get the daily A/B test reports. So I can track not just what I'm working on, but also what my team is doing and see the progress of different experiments. And by the way, the idea here is not to like just ship whenever something looks nice.

It's more just to understand the progress, to give some feedback, ask some questions. But our process to ship an experiment is typically based on hitting some kind of milestone of [00:37:00] time or sample size so that we don't like p-hack or do any or have too many false positives so like my life already orients around running experiments, consuming experiments from the group, from the team.

I think pretty much everyone in the team like already lives and breathes experiments. I don't think that's gonna change. I don't think that any of the AI improvements that are happening change this culture at all. What we discussed about just previously on the reading images example where like we have tried to rate, like even getting images read in the right way can be tricky, and there are a lot of corner cases.

That applies, I would say, broadly to like any LLM problem. In my experience that like even the latest and the greatest LLMs might still have failure modes. And I would say that also after doing this like dozens of times, that's actually very rare that if you just swap, let's say, an LLM, a previous one to a newer one, even if it's much better in theory, that's gonna perform as well as the previous one in all situations.

I always find that there is always [00:38:00] alpha. There is always like a leverage by better using this latest LLM. My perception is that like new LLMs, new technology provides a new ceiling of quality, of potential, but ultimately, you are really responsible for going from wherever you are to closer to the ceiling.

And the way to do so is very complicated, I would say broadly speaking. It's not just about experimentation, but experimentation is gonna remain playing a key role into after you have a actual new system, a new proposal, a new prompt, a new architecture, a new custom model that you have trained, it's really gonna be the way that you make decisions.

It's the basics of statistics, the basic fundamentals of like causal inference are just not gonna change, right? The only way to know if a change that you made to the system is causal, is gonna actually... Is the thing that's impacting the system is through experimentation. I don't think there's any way out of this.

I don't think AI solves that problem. What actually AI offers to you, I think what's-- how it's gonna evolve is that AI is gonna [00:39:00] launch more experiments in your behalf, like we're doing r-already. Claude is the main driver for our experiments now, of launching new experiments. AI is gonna analyze more experiments.

It's gonna dredge through the data. It's gonna look at the summary metrics. It's gonna call you out when there's a problem. It's gonna iterate, and hopefully this becomes more and more autonomous. And as team leads, I wanna just look at the report of here's the hypothesis tested here are the results.

And this can be done by AI eventually or at least most of it. What is the role of humans to play here? It's actually a good question. It's a bit unclear. But I think experimentation maybe is gonna change-- the driver is gonna change from humans to more AI-driven. But it's gonna remain as like a tool that like whatever system, whether human-based or AI-based, need to use to make decisions.

Decision-making is always gonna remain fundamental, and decision-making can only essentially be done through experiments.

Ashley Stirrup: Got it. And so are you saying Claude is like proposing a number of experiments and then you have a human in the loop saying, [00:40:00] "Okay, that's a good one, run that. No, this one's not right," that type of thing?

Pedro Tabacof: We are not quite there yet, I think where Claude is proactively proposing experiments. Right now Claude still requires this tier of a scientist or engineer. But Claude can already like propose the experiment in very precise terms write the whole scaffolding for the experiments do a lot of offline evaluation to make sure the experiment tests what you want to test for real.

So for example, let's say I have a new prompt to write. Instead of writing the prompt myself, as I would do two years ago, what I'm gonna do now with Claude is the following. First, I'm gonna establish some kind of like offline evaluation protocol, so methodology. Maybe I'm gonna use some, a mix of deterministic metrics and the LLM judges metrics.

And once I'm confident that we have a methodology that works, that's aligned with whatever I intend to do, I'm gonna let Claude churn through a lot of prompts and LLMs and use this offline metric to judge whatever is the best one. And then once we have a winner offline, we are gonna have to [00:41:00] put it to the test in a live experiment fashion.

And again, Claude is gonna be the one that's gonna bring this online. And I'm just gonna tell Claude "Hey, now that we have found this winner combination of prompt LLM, let's create a real A/B test with that," and it's gonna do it for me, and I'm only gonna have to review the codes and then monitor the experiments.

But hopefully, as time progresses Claude becomes more proactive in like maybe going from, for example... I think this could be done today already but may-maybe it's too aggressive, but you could, for example, in the future have Claude read like a customer tickets, like customer reports, right? Our own like customer support problems at Fin's our own complaints.

And based on whatever our customers are talking about or complaining about, Claude could create like its own experiments, ship them and monitor them as well. I don't think we are really far off from that reality. That's probably where we're gonna progress. For now, the human remains in the loop steering Claude towards what should be done.

Claude's just executing. We let's say w-we [00:42:00] humans this year, we own the strategy but Claude's owning the tactics. Maybe that's gonna change in the future.

Ashley Stirrup: Yeah. I love how it's clear that you and the team are really thinking about how you just keep scaling experimentation. How can you take AI to the next level? And it's gonna be super interesting to see how quickly AI can take on more things and where we decide, okay, these are the places, AI is not good at and, requires more human intervention, and these are places that we can trust AI more and more over time.

Pedro Tabacof: Yeah. Today, I feel that AI lacks a lot in taste, in just like what is actually important, what should be prioritized. So AI is very good at the low level, like writing code writing prompts even churning through ideas. But knowing exactly what to experiment, that's like really the big challenge.

And ultimately, that's what matters to the business the most, right? You can run a thousand useless experiments, changing like a lot of button colors which might have some small marginal impacts, but sometimes there's one big question you have to answer [00:43:00] as a business, one big thing to experiment on.

And I feel like Claude's not that great at those big questions to be asked. Humans are still like the drivers. So we go a more-- Our roles evolve to be a bit more high level, to understand what the business is doing, what are the needs, like how the system works, and what we can do to improve the system to align better with the business, but then let Claude execute on the challenge.

Ashley Stirrup: Yeah. Yeah, makes total sense. I can't help but thinking, boy, it'd be fun to have you come back on in a year and talk about, like, how much has changed. It's probably gonna... We'll probably be blown away by how much things have changed in just a year.

Pedro Tabacof: Definitely. If you just look at how things were done one year ago, where Claude was not even that popular, right? One year ago, people were still using like mostly some sort of auto-completes or just copy and pasting code arounds, and now we have agents, controlling agents. I think things can evolve in a crazy way in, in one year's time.

Hopefully we are still gonna remain in the loop and with some value in our taste and [00:44:00] strategy still gonna matter. I think that's very likely, but who knows?

Ashley Stirrup: Yeah. No, I'm a big believer that AI is an amplifier for people. So I think we'll still all be employed and just doing amazing things,

Pedro Tabacof: it's yours to dance

Ashley Stirrup: Pedro, thank you so much for joining today's episode. It was just a wealth of information. Really appreciate you coming on the show.

Pedro Tabacof: Yeah, I think this was like a really nice the questions were great. Yeah. I'll be happy to come again in a year's time and see how things have evolved since then.

Ashley Stirrup: You can count on me definitely asking you. Thank you so much

Pedro Tabacof: Cool. Yeah.

About Pedro Tabacof

Pedro Tabacof is Principal Machine Learning Scientist at Fin, formerly Intercom, where he has spent three years running the experiments behind one of the world's most advanced AI customer support agents. He focuses on A/B testing non-deterministic AI at scale, pulling millions of samples in days to turn surprising results into shipped improvements users can trust.

Role
Data Scientist
Industry
Business Tech

Subscribe to the podcast

Takeaways from this conversation

Fin A/B tests everything, even one-character prompt changes, and treats a 20 to 30 percent win rate as a sign of a healthy program.

S1 | E28

A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.

S1 | E28

More conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations.

S1 | E28

You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change helped.

S1 | E28

Faster is not always better. Fin raised latency artificially and positive feedback went up, likely because a small delay makes an AI feel like real work.

S1 | E28

Top takeaways from other favorite conversations

All Takeaways

Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.

S1 | E7

Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.

S1 | E26

A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.

S1 | E22

Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.

S1 | E13

One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards.

S1 | E21

Reposition features around how users actually feel, not how you assume they should feel

S1 | E9
The experimentation edge podcast logo with a picture of host Ashley Stirrup