Back to Podcast
Testing AI
AI-Native Dev
Culture

ServiceNow's Customer Zero approach to AI experimentation

S1 | E45
Sep 29, 2026

Summary

ServiceNow runs its own platform on itself. Ashraf Karim, Senior Vice President of Connected Customer Experiences and Technology, owns everything that happens after a customer buys, and her team is Customer Zero, testing every product and capability internally before it ever reaches a customer.

She joins Ashley Stirrup to unpack what that vantage point teaches you about experimenting with AI. The headline lesson is not a technical one. ServiceNow built agentic skills that worked, and users abandoned them anyway, because a task that ran five to fifteen minutes broke every expectation people had about what a machine should feel like. The fix was not a faster model. It was telling the user what the AI was doing while it did it.

Ashraf also gets into the harder measurement problem underneath all of it: how do you evaluate a nondeterministic system? ServiceNow's answer is an internal LLM judge scored against a golden data set, plus a customer effort score that catches what CSAT reports too late.

📝 Read the full blog post →

Chapters

00:00 What customers actually care about
01:14 Meet Ashraf Karim of ServiceNow
01:39 Owning the entire post-sale experience
02:08 Why ServiceNow runs as its own Customer Zero
03:02 Lessons from Google, PayPal and Verizon
07:13 The agentic skills users kept abandoning
09:31 Latency, expectations and the feedback loop fix
12:31 Users wanted the outcome, not the process
14:41 Designing experiments that can afford to lose
18:06 Using an LLM judge on nondeterministic models
21:41 Customer effort score as a guardrail metric
24:56 Where experimentation at ServiceNow goes next

Notable Quotes

"We are also Customer Zero. And what that means is that we test the products, the features, the capabilities before they become available to customers."

"It's not that technology is not working, technology is doing its job, but it's really around the performance and latency. The customer expectation, the user expectation is different."

"Our users didn't want us to use that process and they just wanted to know what was the outcome of it. Just do the work that you're supposed to do and just tell me what I should do next and what's my next best action."

"There is no bad ideas, there's just ideas. And if you were to look at every single idea from the vantage point of feedback and learning, then you have not lost. You have only to gain."

"I think of CSAT as a lagging indicator, and CES is more specific to that particular job, an intent that we expect that the customer was here for."

Transcript

Ashley Stirrup: Hello and welcome to today's episode. Today I'm excited to have Ashraf Karim, Senior Vice President Connected Customer Experiences and Technology at ServiceNow. Welcome to the show, Ashraf.

Ashraf Karim: Well, thank you. It's great to be here.

Ashley Stirrup: Yeah, I'm excited to have you talk about all the exciting things that ServiceNow is doing. Why don't we kick things off by having you tell us a little bit about your role there?

Ashraf Karim: Yeah, of course. So here at ServiceNow, at the end of the day, if you think about everything that happens after a customer purchases, the entire post-sale experience, that's essentially what my team and I are responsible for. And the way we do this is essentially we look at how do we power ServiceNow on ServiceNow to provide a great customer experience. So we look at how do we enable the agentic workforce, how do we provide an effortless customer experience. And we do this all within ServiceNow. So we are also Customer Zero. And what that means is that we test the products, the features, the capabilities before they become available to customers. And we use the platform itself and these brand new features to actually run ServiceNow altogether. So everything post-sale, as you'd look at fulfillment, servicing, and how do we also grow our customers in general with ServiceNow University, our developer portal, our developer ecosystem. These are all the wonderful things that my team and I oversee. And yeah, we look at productivities, we look at efficiencies, we look at how do we generate more revenue, how do we get time to value. We've got so many different KPIs, but it's all about the post-sale experience and how do we make it effortless and compelling and great for the customer.

Ashley Stirrup: You've definitely got a big job, especially at a company the size of ServiceNow. You've also had some great experience at Google and PayPal. Could you tell us a little bit about that as well?

Ashraf Karim: Yeah, of course. You know, they're great companies. A lot of my work there also was how do we build efficiency and productivity to drive self-service? How do we become more productive with tools and technology? But the primary focus when you look at these companies, it's really consumer-based. So a lot of the products and tools and technologies that we built was to serve, think about B2C, but how do we directly serve those customers? So the volume of customers that we get, the types of technologies we build, like for example PayPal, PayPal system taking it zero to one, right? It's a complete B2C type of experience. So very different in that sense of consumer versus what we do at ServiceNow, which is more B2B. But at the end of the day, it's all about serving the customer. So these companies Google, PayPal, Verizon that I previously have been at. are really consumer focused and ServiceNow is again, we're just trying to build a platform that actually serves the end customers directly as well.

Ashley Stirrup: Yeah. And as we were talking a little bit before the show about just how data driven Google was in particular.

Ashraf Karim: this was about over ten years ago or so, right? So when you look at how data driven Google was, and right prior to that I was at Verizon, again, right? You got big data coming about. We're talking about what, 2004 to 2008. Those years, right? Data was just coming about when we were dealing with with high volume variety of data and we were trying to figure out, hey, how do we actually get value out of the data. So at that time, going from Verizon to Google and where you had data at your fingertips, it was an amazing culture. It was an amazing time. And then a few years later after I switched to PayPal, it was interesting because we were just building the foundational part of how to get access to data, how do we use data to generate value. So Part of the work and the transformation that I did at PayPal was really looking at how do we become a data driven organization and how do we start to provide great customer experiences using that data. So it was a journey to get there. And so one of the learnings here is just going from all these wonderful companies, was that not every company has a really strong data foundation, at least for the type of work. And the post-sale experience that I've typically have led, I have found that to be more of an underinvestment area in those types of companies, but tremendously valuable once you actually start to invest in it and you understand the value it can provide.

Ashley Stirrup: that makes a lot of sense. And I think you brought a lot of that kind of data driven and customer mindset to your role at ServiceNow.

Ashraf Karim: Yeah, a hundred percent. Because at the end of the day, right, we are transforming the way we deliver experiences to our customers and outcomes. Right. So what do customers care about in the world that I live in today? They care about time to value, time to adopt. They want to get their support questions answered as quickly as possible. I also have the privilege of leading ServiceNow University and our developer PDIs, personal developer. instances where customers and developers can come and play with ServiceNow capabilities. So when we look at the opportunity here, right, to provide AI learning, we provide simulation studios, so many different capabilities for customers to learn and adopt and explore our products, there's this big opportunity if we use data the right way, we learn about the customers, we learn about how they're using our platform, what are areas of opportunity. Ultimately we can provide a really great, personalized, effortless and efficient experience to help them get to the answers that they want as quickly as possible and how they're using our platform to deliver the outcomes that they're looking for, also to serve the businesses and customers they serve.

Ashley Stirrup: So maybe, is there an example of an experiment you've run where you've had a lot of learnings?

Ashraf Karim: You know, I'd say learnings are everywhere. It's from your early childhood to where we are today. But one of the things that really stands out is we were building out, I would say, the agentic framework or the agentic workforce for ServiceNow and for our support space. And in that spirit, we were essentially looking at how do we take what we have from customers in the cases that we're getting from customers, and how do we self-serve them and how do we get back to them as quickly as possible with the information they need. So one of the things that we found is that we were deploying and creating these agentic skills, right? We take these skills and we're building them. So say for example, we'd build a skill that would go through thousands of records in a Splunk log file. And it would look at exactly why there was a problem in a customer instance or a customer issue. We would build out a skill to do that. And we would build additional skills as well. That was just one example of a skill. We would probably have another skill that would look through all the records and history that we have within our entire database and our instance. And we would look at, okay, what are similar ways customers, how have we solved customer problems previously. So we would create a skill for that as well. So what was happening is that we were creating so much of these skills that over time these are complex tasks. And these complex tasks can take five to fifteen to twenty minutes, depending on how complex the issue is. Because this is like a high-touch environment, right? So it takes time to resolve and to investigate and explore. And one of the findings that we found was that one of the biggest learnings in this process was that we were finding initially that folks would just abandon or they didn't like using the skill that we launched. But we're like, well, we put so much effort into it, it takes the task away from a human having to do it, you can do something else. But people are getting really frustrated with the experience. So we said, well, is the technology not working? Is that the problem? And then we come to find out really, it's not that technology is not working, technology is doing its job, but it's really around the performance and latency. So the customer expectation, the user expectation is different. So even though a user might have typically spent hours trying to investigate something, asking them to watch a screen do a job for like five, 10, 15 minutes, it's kind of like, why am I doing this? Like this is taking too long. So what we learned in that, some of the learnings that we got out of these types of experiments and these being early adopters of this new technology was that obviously performance and latency plays a big role. But how do we actually also create that feedback loop to the customer, to the user from an experience standpoint? And we tell the customer that this work is happening, and we're constantly sharing what the AI is doing in the background so that they have the patience to continue to wait and know that there is work that's taking place. So that's been some of the learnings along those lines.

Ashley Stirrup: Yeah, boy, that's a really powerful example because so often I find with different guests is with every guest you have to understand their customer and the customer's journey. I often call it a buyer's journey, but it's not always buying, but it's whatever the task is, the job to be done. And then it's understanding which of the features and capabilities that I'm providing is helping on that journey and which one of those maybe it's a good idea, but it's at the wrong step of the journey or it's being maybe in like the example you just gave, not setting the right expectations with the user. And so that to me is where A/B testing can be so powerful, is to test different models, see is the faster one better or is the smarter one better, all those different types of things. And so that's a great example of how you got something out into the market and then kind of tested and iterated on it.

Ashraf Karim: Yeah, for the work that we do, honestly as a technologist, you're always just iterating, right? And software, that's what software is. Software is a process of iteration. You're constantly just improving the product that you just launched to make it better and add more value to it. So yeah, we learn a lot, definitely.

Ashley Stirrup: Yeah. And I think, we're talking before the show, you mentioned that you don't always necessarily want to replicate exactly what the humans were doing as you're building the new AI workflows and that there's things you might cut out and things you might add to it, like we can be much better with AI at doing cross-sell on top of solving a customer problem.

Ashraf Karim: Yeah, there's so much. that's the beauty of where we are. Like if you have the right data infrastructure and you have the ability to use that data to provide that value of cross-sell opportunities, upsell, so much is there for companies to grow in that space, right? And to enrich the portfolio, the customer base, and grow with the customer. It's really around providing that personalized experience at the right moment in time so that the customer can get the value and you can share your insights with them. And one of the things that we learned, just in the spirit of as we were going through a moment of taking what the humans were doing and we were building agentic skills to kind of replace that. One of the things that we learned and this was like an aha moment for us was when we took that same process and we agentified it, and we said, all right team, just use this process. Now you don't need to do it yourself. What we found was that our users didn't want us to use that process and they just wanted to know what was the outcome of it. So it was like, so if we were troubleshooting an issue, they didn't care about the steps that they did for troubleshooting. They're like, well, why don't you just tell me what I need to do next? It's as simple as that. So do the work behind the scenes. And I think this is the power of the platform that we have. And this is the power of the agentic workforce. Just do the work that you're supposed to do and just tell me what I should do next and what's my next best action. Or just guide me. So it's super interesting the evolution of how we actually do servicing and how we actually drive the post-sale experience with an agentic platform. The expectations of the customers are a lot different when you have AI at the center.

Ashley Stirrup: boy, that's a really interesting topic, because I could imagine that you have to build trust to get to that point. That maybe initially it's like, show me all the things you're doing. Let me make sure you didn't skip any steps. And then once you start to trust it, okay, you don't need to show me all this, I know you're doing those steps. Now just tell me what to do. So you could see that being a little bit of a moving target.

Ashraf Karim: it is. It is a moving target. And I think just knowing that, right? This is what it takes to build trust. So it's the process of it. And it also makes great products too, because in that journey of that validation, you're also refining and improving the technology as you go along as well. So there is that really good verification, test and verify type of process that's embedded in it.

Ashley Stirrup: Yeah. So let's say you're working with somebody kind of newer to experimentation and maybe they haven't fully got their head around the idea that you need to be humble and that you often don't know what you don't know until you've run the A/B test. And so you end up with more losers than you do winners. And so therefore, of course, as you're designing an experiment, you want to be thinking through, well, what if this is a loser? How do I design the experiment so I have enough data to really understand what happened so that I can iterate and can learn as much as possible from this experiment. How do you think about that?

Ashraf Karim: Yeah, it's a mindset of how you approach testing in general. Because at the end of the day, we try a bunch of ideas as creative humans, right? It's like we had a bunch of ideas out there. There is no bad ideas, there's just ideas. And if you were to look at every single idea from the vantage point of feedback and learning, then you have not lost. You have only to gain. And I think that is the most important piece. Look, it's really trying to understand what can you learn from this. And there is learnings at every single point in time. And when you bring that learning in, it makes your next experiment better. And ultimately, after enough experiments and trials and tests and learnings, you essentially build a great product and a great solution out there. And that's really how, like one of the things that we're very early adopters actually. of ServiceNow ACR platform, which is autonomous case resolution platform. In order for us to actually go build this and take it to the market, it takes a lot of trial and error. It's a lot of trial and learning. Let me put it that way. So it's like you try with your first agent, you see, hey, did this agent work? And you come to find out, well what? We need to move faster, we need better latency, we need better performance. All right, build out the next skill or a different skill. And you come to find out what I'm proposing, it's similar to what a human might say, but because it's not coming from a human, customers want more validation. So you learn that. So it's kind of like you're constantly trying to put something great and better for the customer. And then you gotta constantly be open to, okay, did you get the lift? Are people using it? What's the customer sentiment? What is the CSAT? So you have all of these different metrics that you have instrumented and you have in mind of saying, okay, this is how I'm gonna know that I delivered an effective technology experience, right? So is the technology working well, it's also the experience layer. You gotta decouple those, so they're very related. So you have to really think about how you build out your experiments to test out different components of the stack. and of the solution you're delivering.

Ashley Stirrup: Yeah. Well, I can't help but think about all the moving parts that you kind of just touched on. And the beauty with AI is you can move faster. The challenge with AI is it's less deterministic. And so you can use the same prompt and the same model and get different answers different times. And so lots of opportunity to iterate, but it's super important to be data driven while you're doing it.

Ashraf Karim: Yeah, it is. And I love that you just mentioned the idea that, right, when you go with LLMs, these frontier models are nondeterministic models. And how do you then evaluate the output that you have there? So one of the things, and we ran into this too, because when you were deterministic, it was very simple. You knew exactly what the next answer was gonna be and you go back and forth. But at the end of the day, you want to create an AI experience that's conversational like and natural like. And in that sense, there is fluidity and it is nondeterministic. So one of the ways we solved this was internally we built what we would call an LLM judge that would essentially take the output of our nondeterministic model, and it would take what was a good output of our golden data set. And it would do a quick comparison and evaluation, say, okay, how close is it to this? How confident am I that I understood the intent and that this output is going to serve the customer the way we intend? And then we look at the output of that LLM judge, it gives a score. And based on the score of 70, 80, 90 percent, depending on the threshold that we're comfortable with, we would then take that output and send it back to the customer and say, okay, I think we have a good proposal for you. But it's really key to still let the customer have complete control. So even though we would use our LLM judge to tell us that our nondeterministic output was good or bad, if it went back to the customer as a proposal, we would still give the customer the option of accepting it or following up with us on it to come back with more questions because they didn't like it, in which case we would, you know. take another look at it or we would just route it to a live agent to look at more closely as well.

Ashley Stirrup: Yeah, makes a lot of sense. Yeah, I think LLM as a judge and AI evals can play a really important role in the kind of, I would almost call it the QA side of like, okay, is the new prompt better than the old prompt? That type of thing. But super interesting there's certain use cases where AI is very clear. For example, GrowthBook customer Typeform, and they added an AI wizard to help with the creation of a survey. And so it's very easy to measure how long does it take somebody to create a survey? Did they publish it? You know, those two metrics kind of tell you did the AI add value or not. But in a customer service example or a blog writing example, things can be more subtle. You might have given the answer to a customer, but did that really resolve their issue or did they leave frustrated? Or did somebody write a blog and then like go rewrite it? That type of thing. And so it requires more thought and creativity on the metrics and like just understanding was this really a good customer experience or not.

Ashraf Karim: a hundred percent. And it's interesting, right? Because in a consumer world your metrics, right? Your engagement metrics. You've got your click through rate, conversion rate, you see the person time on site, all of these wonderful metrics that will come and tell you how well a person is interacting with the content that you provided. Whereas when it comes to the type of work we do, especially in the agentic world, right? the feedback is ultimately from the customer, right? And sometimes it's binary. Like it's kind of like it was a hit or a miss. But I do feel there is a gray area in between. And it may not be a hundred percent perfect, but are you at least satisfied? So for that reason, one of the things that we also did introduce, and I think you touched upon this earlier, like the concept of the jobs to be done, right? Like a customer came to us for a specific intent, a job to be done, how well did we actually serve your need and how effortless was it? So we recently introduced a customer effort score, CES is what we call it. And we use CES to tell us how effective or how effortless was the experience that you just had with our technology. So that's one of the guardrails. It's another way of looking at CSAT, but I think of CSAT as a lagging indicator, and CES is more specific to that particular job, an intent that we expect that the customer was here for. But in addition to this, we also, if you look at the funnel itself, like a conversion funnel for an agentic bot or an agentic platform, you're really looking at, okay, did the person get through, how effective it was? So there's a completion rate where you can say, okay, did the person go through all the way to the very end? They engaged enough, and it's all natural language, right? So we're going back and forth. But at the end of the day, did the customer get all the way through the entire conversation with us and we ended on a positive note with a recommendation in place? Then what we also have to look at is did the person come back in the next 24 to 48 hours or within the next week for the same issue? Or did they come back and they could have come back, but they would have come back for a different issue? So you have to think about the intent being really core to understanding if a customer's user behavior is such that they come back, why did they come back? And then you can say something has been resolved and satisfactorily addressed. That's a little bit of some of the metrics that we've had to look at.

Ashley Stirrup: those are really great examples of they might not be your primary metrics, but really important secondary metrics to understand are you delivering the value or not? I love that whole customer effort score.

Ashraf Karim: Yeah, we have like, there's so many primary metrics, right? Because your primary metrics are usually business driven, right? Hey, how much revenue did you bring in? What was the operational savings? If the case itself was resolved, how much did you save, right? If there was productivity, then we're looking at how much efficiencies did we gain from if you didn't have a human do it and if an agentic capability did it, what was the time savings associated with it? So there's so many different higher level metrics, but at the end of the day, as a technologist, you want to build a great customer experience, you want to build a great technology to enable that. So really, how do you measure a powerful technology and how do you measure how well that technology performs for the customer and how do they feel about it? So I think these are, I guess it would, you'd say guardrail metrics or secondary metrics. But they're super important, informative. Like I consider them as platform metrics and product metrics, because it just really tells you how great is the product and the platform that you just built.

Ashley Stirrup: Yeah, it can be tough to pick a primary metric in a situation like that when you've got a number of key metrics.

Ashraf Karim: Yeah, and it depends on who you're talking to, I guess.

Ashley Stirrup: Right. Well that's true too. So how do you see experimentation evolving at ServiceNow going forward?

Ashraf Karim: You know, there's just so much possibility and so much opportunity because we're constantly innovating, we're building out new technology. So we have to live and breathe experimentation, right? So there is so much for us to do from that vantage point. And I just see us more and more experimenting, learning, and launching more frequently than we've had in the past and moving a lot more faster because technology is moving so fast. So keeping up with that pace and technology. is exciting and that requires us to do a lot more experimentation and user testing from the very beginning. So if anything it's just keeping up with the pace and using this as a feedback loop, using experimentation as a feedback loop constantly to improve our products. So we do a lot of it today. We're just gonna do more and more of it. And there'll be multiple variants out there, right? A, B, C, and D, like let's test it all. So It's pretty exciting from that lens, but lots of possibilities.

Ashley Stirrup: It is. It's pretty impressive. I just think about it's one thing to be building a chat bot for your own website. And it's another to be building a set of customer service tools that thousands of customers are going to use with thousands of different use cases. And what performed well over here might not perform as well over there. And so that's a pretty daunting set of challenges you're taking on, to like how do you optimize for that kind of breadth of customer base.

Ashraf Karim: Yeah, a hundred percent. And we see it all. Like we have our low touch cases, we see our high touch cases, similar platforms, similar technology, but very complex problems. But it's very exciting and it's an exciting time and I think technology, we're at the right time, at the right place to really see its potential within the platform as well.

Ashley Stirrup: it's gotta be a very exciting time at ServiceNow. Just all the value you can add to the whole kind of AI ecosystem and the wave of new products that are getting delivered out there. Yeah. Well, thank you so much for joining the show today. I feel like we learned a lot about the ServiceNow business and just how you're tackling customer service challenges. So thank you so much.

Ashraf Karim: Thank you, Ashley. It's great to connect and it's a great conversation. Thank you.

Ashley Stirrup: Thank you.

Ashraf Karim is Senior Vice President of Connected Customer Experiences and Technology at ServiceNow, where she owns the entire post-sale experience and leads the company's Customer Zero program, running ServiceNow on ServiceNow. Her background spans data-driven customer experience work at Google, PayPal and Verizon. She brings deep expertise in agentic AI adoption and in measuring nondeterministic systems with LLM judges and customer effort scores.

LinkedIn
Ashraf Karim
Role
Exec
Industry
Business Tech

Subscribe to the podcast

Takeaways from this conversation

Nondeterministic output needs its own evaluation layer, so ServiceNow built an LLM judge that scores against a golden data set before anything reaches a customer.

S1 | E45

Users did not want the agentified version of the human process. They wanted the outcome and the next best action, not the steps.

S1 | E45

The fix was a feedback loop that narrates what the AI is doing in the background, which buys the patience a long task needs.

S1 | E45

The agentic skills worked. Users still abandoned them, because a five to fifteen minute wait broke their expectation of what AI should feel like.

S1 | E45

ServiceNow is its own Customer Zero, running the platform on itself and testing every feature internally before customers see it.

S1 | E45

Top takeaways from other favorite conversations

All Takeaways

Prioritize like a pyramid: fix the widest-impact experiences first, then optimize down into smaller cohorts.

Go to S1 | E26
Theme
Scale
Role
Product
Industry
Business Tech
Featured
true

Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged.

Go to S1 | E36
Theme
Culture
Role
Data Scientist
Industry
Media & Gaming
Featured
false

Design every test so it teaches you something whether it wins, loses or ends flat. Losing tests are jet fuel when the learning is built in.

Go to S1 | E43
Theme
A/B Testing
Role
Product
Industry
Consumer Tech
Featured
false

Turn data into narratives with AI to deepen engagement and increase discovery.

Go to S1 | E8
Theme
AI-Native Dev
Role
Exec
Industry
Consumer Tech
Featured
false

Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.

Go to S1 | E1
Theme
Culture
Role
Exec
Industry
Business Tech
Featured
true

Grubhub tests big bets by releasing to internal users or small cohorts first, so it can read business impact without disrupting core revenue and order metrics.

Go to S | E
Theme
Scale
Role
Product
Industry
Marketplace
Featured
false

Scaling past low hundreds of experiments per year is a capabilities problem before it's an AI problem — Home Depot is moving from client-side to server-side testing so winners release quickly, end to end.

Go to S1 | E21
Theme
Scale
Role
Data Scientist
Industry
Retail
Featured
false

Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.

Go to S1 | E10
Theme
Culture
Role
Exec
Industry
Retail
Featured
true

Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.

Go to S1 | E26
Theme
Growth
Role
Product
Industry
Business Tech
Featured
true
The experimentation edge podcast logo with a picture of host Ashley Stirrup