Why Twilio ships on signals instead of significance
Summary
Wanli Lau spent eleven years at Expedia Group, most recently running consumer product analytics on a program that shipped more than 1,500 A/B tests a year. Five months ago she moved to Twilio to lead global R&D analytics. In this episode she walks Ashley Stirrup through what carried over and what did not.
The centerpiece is a test that looked like a win. At Expedia, an account sign-up takeover on the front page drove conversion sharply up, and quietly harmed return rate and engagement across every geography. That result is where Wanli's insistence on guardrail metrics and a trade-off decision framework comes from. She also explains the parts of B2B experimentation that have no B2C equivalent: choosing between user-level and account-level randomization, the interference risk when two colleagues at the same company see different pricing, and the many-to-many problem of one developer belonging to several accounts. And because Twilio's customer counts are lower than a consumer business, her teams often ship on signals and team conviction rather than waiting for statistical significance.
Chapters
00:00 Cold open
01:02 Welcome Wanli Lau of Twilio
01:23 Inside Wanli's role at Twilio
04:03 Eleven years and 1,500 tests a year at Expedia
04:58 Why B2B experimentation differs from B2C
05:33 Cross-functional collaboration and good hypotheses
06:39 Teaching teams to experiment rigorously
07:48 The Expedia sign-up takeover that won on conversion
11:09 Guardrail metrics and the trade-off framework
14:52 Designing experiments to maximize learning
17:46 Diagnosing where a feature failed
19:10 North Star metrics and account-level randomization at Twilio
21:40 Learning from signals instead of significance
22:11 Where experimentation goes next at Twilio
Notable Quotes
"The cross-functional collaboration is really critical to ensure that experimentation is having the good hypotheses and then we have the good success metrics and primary and secondary metrics and guardrail metrics that are set up."
"The conversion was even up, but it was negatively impacting the customer return behavior and their engagement. So it was harming the guardrail metrics and that's why we needed to do more UX research and iterative design."
"Not every team will be driving conversion rate as their primary metric, but conversion rate should be always do no harm as a guardrail."
"If the product managers didn't decide in the very beginning what are the number one, number two primary metrics, when the test readout time comes, then you might debate like, okay, there's a little engagement uplift, or can we call this test as a winner?"
"Twilio's customers count is not as high as other consumer facing companies. So in many implementation we are not necessarily driving for the significance. We are learning from the signals. And if it's good enough and the team has a really strong conviction, then we would ship the test design."
Transcript
Ashley Stirrup: Hello and welcome to today's episode. Today we're excited to have Wanli Lau, Senior Director, Data Science and Analytics at Twilio. Welcome to the show, Wanli.
Wanli Lau: Thank you, Ashley. Happy to be here.
Ashley Stirrup: Maybe we could start things off by you telling us a little bit about your role at Twilio.
Wanli Lau: Yeah, so here at Twilio, I am five months in. I am leading the global R&D analytics, and my team is called Data Science and Analytics. Now we are actually expanding to the entire analytics and data end-to-end with the data platform as well. And Twilio is a company that is a leading organization and enterprise to provide an API for the developers and businesses to embed the APIs into communications. Those include many different formats, such as voice, messaging, that will include SMS, WhatsApp, and email as well as video. And we also have the customer data platform like Segment and also the Flex platform for customer support. So you can see that Twilio is really pushing for that communication platform as a service. And this year we are launching the new product called Conversations Products with four or five different specific suites to leverage AI to help our consumers to really make the communication a lot more seamless. And my role at Twilio is to provide data, some implementation, understanding the customers' adoption journeys, and providing the analytical and business recommendations for our product innovation.
Ashley Stirrup: Yeah. I'd imagine with a business like yours, I mean Twilio's been very successful, went public I guess about a decade ago now.
Wanli Lau: That's right. We just celebrated a ten year anniversary.
Ashley Stirrup: I remember it like it was yesterday. I was at a company that was thinking about going public and we watched Twilio's IPO very carefully. It was a well done IPO at the time. And so I would imagine with a business like yours, there's just so much information you have on your customers and how they use your products and all that. I would imagine there's a lot of different demands on you and your time. Is that true?
Wanli Lau: Right, and since I'm only five months in, I have seen many different things, but there are still a ton of learning for me to deep dive. But implementation is really important, especially in our Twilio.com website and also our Twilio console. That is the console where the users will log in to explore our products and try out our APIs. The customers who are interested to use the product, they have the trial credits. We as an analytics team try to understand the customer journey there. How do we enable the onboarding and compliance, regulatory demand in a more seamless way so then people can get their phone numbers and get their voice running and started to try out the products.
Ashley Stirrup: Yeah. And before Twilio you did a lot of A/B testing at Expedia, right?
Wanli Lau: Yeah, that's right. I spent 11 years at Expedia Group. My latest, I would say five years, is running the consumer product analytics team driving the growth initiatives. The most notable ones including acquisition and landing pages, the growth acquisition, as well as the app experiences, personalization, loyalty and communications and cross-sale. So there are many, many implementations that my team has run with the product and engineering and design. All together at Expedia during my time there was more than fifteen hundred tests run during the year. So that was really amazing time where we had tons of learnings and also at the same time building the very fantastic implementation platform.
Ashley Stirrup: Yeah. And I would imagine A/B testing is quite different at Twilio with more of a B2B business model than it was at Expedia.
Wanli Lau: That's right. So at Twilio, because it's B2B businesses, we have our websites, we have our consoles, but in terms of the sample sizes and the customer profiles, they are quite different compared to Expedia, which is mostly consumer driven B2C businesses. But Expedia also has a B2B business. I wanted to call that out, but my focus at that time was at consumer product.
Ashley Stirrup: Got it. Sounds good. And so are you working with many different product teams? When it comes to A/B testing, do you find that you have to kind of coordinate across lots of teams, share ideas, share learnings across those groups?
Wanli Lau: Correct. And I think that is the nature in every company that I go to. The cross-functional collaboration is really critical to ensure that experimentation is having the good hypotheses and then we have the good success metrics and primary and secondary metrics and guardrail metrics that are set up. And the data science analytics function is to define the sample sizes, test durations, and ensuring the analytical readout is sound and robust so that we can have the quick decision learning as well as the next step in terms of iterations. So we are working with not just product but engineering in terms of the implementation of the design and obviously designers to brainstorm what should be the right test that we can run to increase our top line and also primary metrics.
Ashley Stirrup: And do you feel like part of your role is kind of teaching people how to run experiments in a rigorous way? I would imagine there's a lot of novices out there who don't realize that, if I do the assignment this way and my key metric is that, I'm gonna get myself into trouble, that kind of thing.
Wanli Lau: I want to shout out that Twilio's product development organization is really, really savvy by leveraging Amplitude capabilities, which is our analytical online platform, to not just reading out the customer engagement and click-through rate, but also running implementation. So the data science and analytics role here is more enablement and also providing the consultation. A lot of the metrics readout will be coming out of the Amplitude interface as long as our data instrumentation and event logging are set up properly. So really wanted to give a shout out to the product managers who can really self-serve and really having that flywheel running a lot more faster.
Ashley Stirrup: That's great. That's a super powerful thing to have. Can you tell us about a time when you ran an experiment where you had a lot of learnings? Maybe something you did from your Expedia days?
Wanli Lau: Yeah, so I think maybe Expedia days will be a better learning given I have spent eleven years there. I would like to call out the experiments that were in the storefront where we are trying to increase the customer logging and then so people can sign up for the One Key program, which is the Expedia loyalty program. It is pretty common for any consumer products to really ensure that the onboarding process is really smooth for the customers. But this is like maybe six, seven years ago. There were a lot of focus on conversions and less focus on the long-term journey and customer engagement. And given conversion was the primary metric and pretty much the North Star for every implementation, at that time the very first version of the A/B test was the account sign up takeover in the front page. And we drove really, really high conversion rate because of so, but it's so intrusive. And then it was harming on other secondary metrics such as customers' return rate and engagement with other modules on the website and also our app. And the impact across the different geographies and global sites were varied, but consistently it was all pretty negative. So that was really profound learning. Maybe nowadays it is like the no-brainer, it shouldn't be such an intrusive experience for the customer. But back at the time, because of the dedicated focus on conversion optimization, there was like the quick learning that we got. And after that we did a lot more user research and did a few different versions of the experiments and then eventually we were able to get to a much better sign-on experiences.
Ashley Stirrup: And so basically you really kind of took over the home page to try to drive people to actually sign up. And so you were able to get more signups, but that didn't necessarily lead to more business over time.
Wanli Lau: Right, because the conversion was even up, but it was negatively impacting the customer return behavior and their engagement. So it was harming the guardrail metrics and that's why we needed to do more UX research and iterative design.
Ashley Stirrup: Got it. And did that all show up on the immediate testing or is that something you kind of learned over time?
Wanli Lau: So because it was not immediately detected that it was harming for more long-term metric, so it was not the immediate action, but then soon later, we got the customer complaints that this was not really helping them in terms of discovering the different trips and destination and dreaming. So product and engineering and later on analytics got on board to try to fix this problem. But in the beginning it was a successful story, right? Because this conversion was really successful.
Ashley Stirrup: Right. It just really reinforces that not only do you need to know your customer, but you need to know like the job to be done or whatever you want to call it, the buyer's journey that they're on. And the first thing you want to do, I'm guessing, is get them hooked on, Expedia has a great hotel or trip for me to go on, then I'm ready to go sign up. And that if you do things in the opposite order, they might sign up, but they haven't actually had that aha moment and then you lose them longer term. Is that right?
Wanli Lau: Exactly. And I think this is going back to the good hypotheses and what is good looking like. And I think the good hypothesis needed to have the conversion as a focus, but we needed to incorporate the guardrail metrics and guardrail thinking in test design. So that is a profound learning, but does not mean that it was a failure, it was a learning, right? And because of so we also extended that learning into setting up the escalation path for the trade-off decisions. And the trade-off decisions can be found in many different types of tests. Like for example, we have a bunch of sponsored ads in our search results page. That is driving revenue, but at the same time it might be negatively impacting conversion. And earlier we were talking about a sign up design that is like more content engagement as well that might be having the different kind of directions compared to the conversion rate increase. So the trade-off decision framework was set up during that time as well.
Ashley Stirrup: Yeah, that's so important. It was interesting. We had the Philadelphia choir on a webinar and they had very different teams. They had a marketing team, they had a team that sold ads, and they had the product team and over time they become a bit siloed and experiments helped them come together to define what success actually looked like. And you know before you kind of had the ads people just wanting to sell more ads and the product people like I don't want too many ads on my page, I want a great user experience. And so helping them create a joint definition of success was incredibly valuable for creating more alignment across the teams.
Wanli Lau: Exactly. And then the other thing that we did at that time, which also continue to be a very big thing, was to understand the customer journey map. All the way from where customers are coming from. That is the discovery and the user acquisition to engagement to conversion to in-trip and post-trip, and how do we also incorporate the customer support in that entire funnel map. And because of so, we have that better idea of that entire end to end journey. When we are thinking about the tests and hypotheses and when we are sitting together to ideate and doing the roadmap for what will be the next things that we wanted to focus on. Each of the subteams will have their focus areas to drive for their key metrics, because not every team will be driving conversion rate as their primary metric, but conversion rate should be always do no harm as a guardrail. And from there, we can really create that thread needle across all the teams to really understand holistically with more than a thousand A/B tests run during the year, how do we move the North Star for Expedia's customer businesses?
Ashley Stirrup: Yeah, I love that. So it's kind of understanding the buyer's journey really leads into my next question, which was like let's say you're working with somebody new at experimentation, they've got a feature they're just so excited about, they just know it's gonna be a winner. So all they're thinking about is testing their winner and not, well, what if this loses? What other information am I gonna want? How can I learn as much as possible? If this is a loser, how can I make sure I'm learning as much as possible? How do you help guide people on how they set up experiments in order to maximize...
Wanli Lau: Yeah, so first we wanted to go back to the customer first kind of mentality. What is the good customer behavior, right? And here it can be mitigating the customer's cognitive friction, it can be maximizing the conversion and eventually revenue. It can be focusing on more long term value for the customer to come back and engage and really having that flywheel. And this area is usually loyalty's play. So once we identify, for each of the sub-teams, what will be the ultimate goals that they are driving? For the questions that you're asking about, the new person joining the team, how can the analytics team help that person to learn about the experimentation process? So we have the experimentation playbook at the time set up at Expedia. And we put together the guide for everyone where we not only look at the entire experimentation workflow, from design to build to capture to analyze and read out, and then how do we learn from that, right? And then for each step of the implementation workflow, we have a specific checklist that anyone who's involved in the experimentation test will need to be doing. And then we have the experimentation review set up. So then people will review every two weeks what will be the new test that we are going to launch, what are the learnings from the last two weeks, and then we have a good documentation that are eventually all logged into our experimentation platform. And earlier I talked about the success primary metrics and secondary metrics and guardrail metrics. The platform would need to configure all those metrics so the dashboard itself would show up those metrics for anyone to look at. But we also really emphasize on no cherry-picking or p-hacking kind of practice, because that is quite common among product managers who are really enthusiastic about what is the outcome of their test. So we have also the playbook or the guidance for that.
Ashley Stirrup: Yeah. Playbooks make a ton of sense. I think the other piece of it is just like, okay, I built this feature. Let's say I thought it was going to create more engagement that would eventually lead to more sales. How do you determine where in the process this feature failed? Like maybe you got the engagement but you didn't get the revenue, or maybe you didn't get the engagement and you know what parts of the user behavior changed and which ones didn't, so that you can get inkling from that to then say, maybe we should iterate on it.
Wanli Lau: Yeah, and because we are able to incorporate more than two or three metrics into the platform, there was another side of problem that came up later on where product managers might say, I wanted to measure 10 metrics for this one test, and then we will run into the comparison problem because if the product managers didn't decide in the very beginning what are the number one, number two primary metrics, when the test readout time comes, then you might debate like, okay, there's a little engagement uplift, or can we call this test as a winner? So in the playbook, we also called out that the success primary and secondary metrics and guardrail metrics needed to be called out and established in the very beginning.
Ashley Stirrup: Yes, yes. Clarity on the hypothesis is so important. So speaking of which, at Twilio, how do you think of North Star metrics? Like are there certain things that you're trying to get the whole team to work on improving?
Wanli Lau: Yeah, so at Twilio there are nuances compared to the Expedia or consumer product companies, because we are working with developers and also working with the companies who are using our APIs, right? So in terms of our implementation strategy, it will be based on user level randomization or based on account level randomization. And for these two types of randomization unit, it will be really important to call out when to use which method and what will be the randomization key and what will be the risk. So, for example, for the user level randomization, it will be mostly used in the UI and UX tweaks. For example, we have different buttons to increase the adoption of the API keys and also individual onboarding flow. So the bucketing is happening at a user ID level. But the risk here is it can be the contagion or interference, right? Because a user only representing one company, I come into one console, I see one price tag. But next time when my colleague also coming from the same company sees a different price tag, but for the same purpose. And so both of us might get confused because we are working in the same company and while we are seeing a different price tag or price determination for Twilio's product. So when we are thinking about the test setup, there will be a risk that we need to call out in advance. So that's one. And then the second thing that is kind of unique at Twilio is about the many-to-many complexity. So one developer can belong to multiple different accounts. So it can be their personal side project and one account representing the enterprise employer. And then from the account perspective, you can have one account like, using Nike as example, Nike as an account, but it contains many users as a developer employee. So when we are dealing with that many to many relationship, we needed to really understand the interference risk that I called out earlier.
Ashley Stirrup: Yeah. Boy, that certainly adds a level of complexity to it all. It's super interesting.
Wanli Lau: Exactly. And also Twilio's customers count is not as high as other consumer facing companies. So in many implementation we are not necessarily driving for the significance. We are learning from the signals. And if it's good enough and the team has a really strong conviction, then we would ship the test design.
Ashley Stirrup: Yeah, makes a lot of sense. So how do you see experimentation evolving at Twilio?
Wanli Lau: So earlier I talked about Amplitude, and I also talked about more data warehouse type of analytics that will join the web analytics information with the financial data and with the customer account data. I think there are maybe two things that I would call out for Twilio's implementation practices. Number one is to double down on the velocity of the implementation using Amplitude. But the second thing is that since we have a lot more richer additional data that can be joined with the web analytics data streams, right? How can we really make that part of the joins a lot more seamless? And right now, earlier I mentioned about data platform that my team is also taking on. So there will be the potential for more learnings for our current implementation. And at the end, it's trying to understand the customer journey and how do we reduce the frictions. If we can find that there's a specific step down of the funnels, then we can have the corresponding reactions and product design to address the customer's pain point.
Ashley Stirrup: Got it. So enable the team to run more experiments, but also kind of continue to enhance your data platform so it allows people to run experiments more easily and get better information from each experiment. Is that a good way to summarize it?
Wanli Lau: Yeah, exactly.
Ashley Stirrup: Well, Wanli, thank you so much for coming on the show today. I feel like we got a great glimpse into two amazing companies, Twilio and Expedia, and I think we learned a lot. So thank you so much.
Wanli Lau: Thank you very much, Ashley. Thank you for this opportunity.
Takeaways from this conversation

When customer counts are low, significance is often out of reach. Learning from signals and shipping on team conviction beats waiting for a number that will never arrive.

Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric.

In B2B the randomization unit is a design decision. User-level bucketing risks two colleagues at one account seeing different prices, and account-level avoids that but costs sample size.

Name the primary, secondary and guardrail metrics before the test runs, not at readout. Deciding afterwards turns a result into a debate.

A test that wins on the primary metric can still be a loss. Expedia's sign-up takeover raised conversion and damaged return rate and engagement at the same time.
Resources
Top takeaways from other favorite conversations

Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.
.avif)
Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.

Build capabilities, not just tests—use experiments to unlock platform features (e.g., metering, paywalls).

You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change helped.

Treat engagement carefully. For a bank, more time in the app isn't a win; trust, fast task completion, and healthy repeat engagement are.

Align on the North Star before building, whether it is revenue, engagement, NPS, or fewer support tickets, and set guardrails so an engagement feature can never quietly drag sales down.

A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control.

