ServiceNow's Customer Zero approach to AI experimentation

Guest: Ashraf Karim, Sr VP, Connected Customer Experiences & Technology, ServiceNow. Host: Ashley Stirrup. Show: The Experimentation Edge. Publishing: September 29, 2026
Most companies test new software on customers. ServiceNow tests it on itself first.
Ashraf Karim is Senior Vice President of Connected Customer Experiences and Technology at ServiceNow, and her remit is everything that happens after a customer buys: fulfillment, servicing, support, ServiceNow University, the developer portal and the developer ecosystem. Her team runs all of it on the ServiceNow platform, and it runs the newest features before anyone outside the company can. "We are also Customer Zero," she told Ashley Stirrup on The Experimentation Edge. "And what that means is that we test the products, the features, the capabilities before they become available to customers."
That vantage point has given her an unusually clear view of what breaks when you put AI agents in front of real users. The lessons are less about models than about people.
A data foundation is not a given
Before ServiceNow, Ashraf spent time at Verizon, Google and PayPal, and the contrast between them shaped how she thinks about experimentation. At Verizon in the mid-2000s, big data was just arriving and teams were still working out how to get value from high-volume, high-variety data. Google, a few years later, was the opposite: "you had data at your fingertips, it was an amazing culture." PayPal was somewhere in between, and part of her work there was building the foundations for a data-driven organization from the ground up.
The lesson she carried forward is blunt. "Not every company has a really strong data foundation," she said, and the post-sale experience she has typically led "is more of an underinvestment area." It is also, in her experience, "tremendously valuable once you actually start to invest in it." That matters for anyone trying to run experiments on customer support, onboarding or adoption: the instrumentation usually has to be built before the first test can be read.
What customers actually care about
Ashraf frames every decision around a short list. "They care about time to value, time to adopt. They want to get their support questions answered as quickly as possible." With ServiceNow University, developer instances, AI learning paths and simulation studios all in her portfolio, the opportunity she describes is the same one every experimentation program chases: use the data to learn how customers actually use the platform, find the gaps, and deliver "a really great, personalized, effortless and efficient experience."
Those metrics, time to value and time to adopt, turn out to be the ones that exposed the biggest surprise in her AI work.
The agentic skills that worked, and that users abandoned anyway
Asked for an experiment with a lot of learning, Ashraf went to the early days of building ServiceNow's agentic workforce for support. Her team was decomposing what human support engineers did into agentic skills. One skill would read through thousands of records in a Splunk log to find why a customer instance was failing. Another would search the company's full case history for how a similar problem had been solved before. Each one took a real, tedious task off a human's plate.
Then usage data came in. "We were finding initially that folks would just abandon or they didn't like using the skill that we launched," she said. The team's first instinct was to look for a bug. There was none. "It's not that technology is not working, technology is doing its job, but it's really around the performance and latency."
The skills were doing complex, high-touch investigations that took five, ten, fifteen minutes to complete. A human might once have spent hours on the same task, but asking that person to sit and watch a screen for fifteen minutes broke every expectation they had about how software behaves. "The customer expectation, the user expectation is different," Ashraf said.
The fix was not a faster model. It was a feedback loop. "We tell the customer that this work is happening, and we're constantly sharing what the AI is doing in the background so that they have the patience to continue to wait and know that there is work that's taking place."
Ashley drew the general lesson: a good capability at the wrong step of the journey, or without the right expectations set, reads as a failure. It is exactly the kind of thing an A/B test can surface, whether the variable is a faster model, a smarter one, or simply a progress indicator.
Users wanted the outcome, not the process
A second discovery came from the same project. When the team took the human troubleshooting process, agentified it step for step, and handed it to users, the users pushed back. "Our users didn't want us to use that process and they just wanted to know what was the outcome of it," Ashraf said. "Just do the work that you're supposed to do and just tell me what I should do next and what's my next best action."
Ashley pointed out the tension underneath that: early on, users may want to see every step so they can trust the system, and only later want it hidden. Ashraf agreed it is a moving target, and one she does not mind. "This is what it takes to build trust. And it also makes great products too, because in that journey of that validation, you're also refining and improving the technology as you go along."
There are no bad ideas, only learnings
Asked how she coaches people new to experimentation, who tend to design tests as if they will win, Ashraf described a mindset rather than a template. "There is no bad ideas, there's just ideas. And if you were to look at every single idea from the vantage point of feedback and learning, then you have not lost. You have only to gain."
That is how ServiceNow became an early adopter of its own Autonomous Case Resolution platform. "It's a lot of trial and learning," she said. Try an agent, find it needs better latency, build the next skill, discover that a recommendation which sounds identical to what a human would say gets less trust because a human did not say it, learn from that too. Each round was instrumented against lift, usage, sentiment and CSAT. Her one structural piece of advice: decouple the technology layer from the experience layer in your test design, because "they're very related" and you need to know which one moved.
Judging a nondeterministic system
The hardest measurement problem in the conversation is one every team shipping LLM features now faces. Traditional software was deterministic; you always knew what the next answer would be. Frontier models are not, and a conversational experience is supposed to be fluid. So how do you evaluate the output?
ServiceNow's answer is an internal LLM judge. It takes the output of the nondeterministic model, compares it against a golden data set of known-good answers, and scores how confident it is that the intent was understood and the response will serve the customer. Above a threshold the team is comfortable with, 70, 80 or 90 percent depending on the use case, the proposal goes to the customer.
Two guardrails sit around it. The customer always keeps control: they can accept the proposal, come back with more questions, or be routed to a live agent. And Ashley noted where the judge's authority ends. LLM-as-judge and evals are strong at the QA question of whether a new prompt beats an old one; they cannot tell you whether a customer left satisfied or frustrated. That still takes experimentation on real users, with metrics designed for the job.
Customer effort score as the guardrail
That is where Ashraf's team added a metric most experimentation programs do not have. Consumer teams live on click-through, conversion and time on site. In agentic support, the feedback is often binary, a hit or a miss, and she believes there is a gray area worth measuring. So the team introduced a customer effort score, CES, tied to the specific job the customer came to do.
"I think of CSAT as a lagging indicator, and CES is more specific to that particular job, an intent that we expect that the customer was here for," she said. Alongside it sit completion rate through the conversation funnel and a return check: did the same customer come back within 24 to 48 hours, or within a week, for the same issue?
None of these are the primary metrics leadership reads first. Those are revenue, operational savings, and time saved when an agent handles a case instead of a person. Ashraf calls the rest platform and product metrics, and argues they are what tell you whether the technology is actually good. "You'd say guardrail metrics or secondary metrics. But they're super important, informative."
Where it goes next
Ashraf expects the pace to increase, not settle. "We have to live and breathe experimentation," she said, describing a future of more frequent launches, more variants tested at once, and experimentation used as a permanent feedback loop rather than a launch gate. Ashley underlined the scale of that: building a support agent for thousands of customers with thousands of different use cases means a result that holds for one segment may not hold for another.
The through line of the conversation is that the AI worked in almost every story Ashraf told. What decided adoption was latency, expectations, trust and effort. Those are experimentation questions, and Customer Zero is where ServiceNow answers them first.
Related articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics — free.




