
Notable Quotes
"The cross-functional collaboration is really critical to ensure that experimentation is having the good hypotheses and then we have the good success metrics and primary and secondary metrics and guardrail metrics that are set up."
"The conversion was even up, but it was negatively impacting the customer return behavior and their engagement. So it was harming the guardrail metrics and that's why we needed to do more UX research and iterative design."
"Not every team will be driving conversion rate as their primary metric, but conversion rate should be always do no harm as a guardrail."
"If the product managers didn't decide in the very beginning what are the number one, number two primary metrics, when the test readout time comes, then you might debate like, okay, there's a little engagement uplift, or can we call this test as a winner?"
"Twilio's customers count is not as high as other consumer facing companies. So in many implementation we are not necessarily driving for the significance. We are learning from the signals. And if it's good enough and the team has a really strong conviction, then we would ship the test design."
Takeaways

When customer counts are low, significance is often out of reach. Learning from signals and shipping on team conviction beats waiting for a number that will never arrive.

Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric.

In B2B the randomization unit is a design decision. User-level bucketing risks two colleagues at one account seeing different prices, and account-level avoids that but costs sample size.

