Why Twilio ships on signals instead of significance

Guest: Wanli Lau, Sr Director, Data Science & Analytics, Twilio. Host: Ashley Stirrup. Show: The Experimentation Edge. Publishing: September 24, 2026
Wanli Lau has run experimentation programs on both sides of a divide most practitioners only ever see one side of. She spent eleven years at Expedia Group, the last five running consumer product analytics on a program that shipped more than 1,500 A/B tests a year across acquisition, landing pages, app experiences, personalization, loyalty and cross-sell. Five months ago she moved to Twilio to lead global R&D analytics, where the customers are developers and enterprises wiring APIs into their own products, and where the sample sizes are a fraction of what a travel marketplace produces in a week.
The interesting part of the conversation is not the contrast itself. It is which lessons survived the move and which ones had to be rebuilt from scratch.
The test that won and lost at the same time
Ashley asked for an experiment that produced a lot of learning. Wanli went back six or seven years, to Expedia's storefront.
The goal was ordinary: get more customers signed in, so more of them would enrol in One Key, the loyalty program. The context is what made it go wrong. "There were a lot of focus on conversions and less focus on the long-term journey and customer engagement," she said. "Given conversion was the primary metric and pretty much the North Star for every implementation, at that time the very first version of the A/B test was the account sign up takeover in the front page."
A takeover is exactly what it sounds like. The sign-up prompt took the homepage.
It worked. "We drove really, really high conversion rate because of so." Then the rest of the picture arrived. The takeover was intrusive enough to damage the metrics nobody had designated as the ones that mattered: return rate, and engagement with other modules on the site and in the app. The effect varied by geography and global site, but the direction did not. "Consistently it was all pretty negative."
Ashley summarised it in the plainest possible terms: "You were able to get more signups, but that didn't necessarily lead to more business over time." Wanli agreed, and named the thing that had been missing. "It was harming the guardrail metrics, and that's why we needed to do more UX research and iterative design."
What makes the story useful is the delay. The damage did not show up in the readout. "It was not immediately detected that it was harming for more long-term metric, so it was not the immediate action." What surfaced it was customer complaints: people were no longer discovering trips and destinations, no longer dreaming. Product, engineering and then analytics were pulled in to fix a test that had already been filed as a win. "In the beginning it was a successful story. Because this conversion was really successful."
Her framing of the outcome is worth stealing wholesale: "So that is a profound learning, but it does not mean that it was a failure. It was a learning."
What came out of it was not a rule against takeovers. It was infrastructure. Guardrail metrics became part of test design rather than a post-mortem tool. And the team built an escalation path for trade-off decisions, because once you admit that two metrics can move in opposite directions, somebody has to decide which one wins. Wanli's examples are the ones every marketplace runs into: sponsored ads in search results that drive revenue while suppressing conversion, and sign-up designs that trade immediate conversion for content engagement. "The trade-off decision framework was set up during that time as well."
One piece of that framework deserves its own line. Not every team at Expedia was measured on conversion. But "conversion rate should always be do no harm as a guardrail." A metric you do not own can still be one you are not allowed to break.
Mapping the journey so the metrics stop fighting
The other durable artifact from that period was a customer journey map, built end to end: discovery and user acquisition, engagement, conversion, in-trip, post-trip, and customer support folded into the same funnel.
The point was not the diagram. It was that with more than a thousand A/B tests running in a year, subteams each optimising their own key metric, the only way to know whether the North Star was actually moving was to see where each team's work sat in one shared picture. "From there, we can really create that thread needle across all the teams."
She is equally practical about how an individual contributor learns any of this. At Expedia the team wrote an experimentation playbook covering the whole workflow, design to build to capture to analyze to read out, with a checklist at every step and a review every two weeks where new tests were proposed and the last fortnight's learnings were walked through. The playbook also carried an explicit warning about cherry-picking and p-hacking, "because that is quite common among product managers who are really enthusiastic about what is the outcome of their test."
And it addressed the failure mode that arrives the moment your platform can track more than three metrics: a product manager asks for ten. "If the product managers didn't decide in the very beginning what are the number one, number two primary metrics, when the test readout time comes, then you might debate, okay, there's a little engagement uplift, or can we call this test as a winner?" The fix is unglamorous and total. Primary, secondary and guardrail metrics get named before the test runs.
What B2B breaks
At Twilio, the hard problems are different, and they start before a single metric is chosen.
The first is the randomization unit. Twilio's customers are companies; its users are the developers inside them. So a test can bucket at the user level or the account level, and Wanli is emphatic that this is a decision to be made deliberately, "to call out when to use which method and what will be the randomization key and what will be the risk."
User-level randomization is the natural fit for UI and UX work: different buttons to lift adoption of API keys, changes to an individual onboarding flow. But it carries an interference risk that does not exist in consumer testing. "A user only representing one company, I come into one console, I see one price tag. But next time when my colleague also coming from the same company sees a different price tag, but for the same purpose." Two people at one customer, looking at two different prices for the same product. That is not a noisy measurement. That is a customer conversation.
The second problem has no B2C analogue at all. "One developer can belong to multiple different accounts." A single person might have a personal side project and an enterprise employer in the same login. Meanwhile an account, Nike in her example, contains many developer employees. The relationship is many-to-many in both directions, and interference risk has to be reasoned about across both.
Then there is the constraint that shapes everything downstream: Twilio simply has fewer customers than a consumer business. Which brings the conversation to the line that gives this episode its title.
"In many implementations we are not necessarily driving for the significance. We are learning from the signals. And if it's good enough and the team has a really strong conviction, then we would ship the test design."
It is worth being precise about what that is and is not. It is not a licence to skip rigor. The same person saying it built the playbook that bans p-hacking and insists primary metrics get named up front. It is an acknowledgement that in a population where significance is frequently unreachable, waiting for it is not caution. It is paralysis with extra steps. The discipline moves from the threshold to everything around it: a clear hypothesis, named metrics, honest guardrails, and a team willing to say out loud that it is shipping on a signal.
Where the leverage is now
Asked how experimentation evolves at Twilio, Wanli named two things, and neither was a new statistical method.
The first is velocity, specifically doubling down on Amplitude, which Twilio's product organization already uses to run tests and read out engagement and click-through themselves. She is notably generous about this. Her function's job there is "more enablement and also providing the consultation," and she went out of her way to credit "the product managers who can really self-serve and really having that flywheel running a lot more faster." A lot of analytics leaders describe self-service as a risk to be managed. She describes it as the flywheel.
The second is joining data that currently sits apart: web analytics streams alongside financial data and customer account data, which is why her team is taking on the data platform as well as analytics. More joined data means more that can be learned from the same test.
And the purpose underneath both is unchanged from the Expedia storefront. "At the end, it's trying to understand the customer journey and how do we reduce the frictions. If we can find that there's a specific step down of the funnels, then we can have the corresponding reactions and product design to address the customer's pain point."
The tooling changed. The sample sizes changed. The question did not.
Related articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics — free.




