The Edge Podcast

How Grubhub tests big product bets before they ship

How Grubhub tests big product bets before they ship

Most conversations about experimentation start at the A/B test. Michal Lenik's starts much earlier.

Michal leads the design teams behind Grubhub's B2B2C experiences: the three-sided campus marketplace, the corporate consumer and corporate enterprise experiences, corporate partnerships, and all of the tooling merchants use. Before that she led design for Grubhub's post-purchase experience, which covers everything after an order is placed, from customer care to the fulfillment app drivers use.

Her admission early in the conversation will surprise anyone who assumes a consumer marketplace runs a constant stream of split tests. Grubhub does not do much classic A/B testing. What it does instead is test earlier, test smaller, and put real options in front of real users before a big bet reaches millions of people.

Design as a business lever, not a screen factory

Michal's teams experiment with products, but they also experiment with process. The question they have been testing is where design belongs.

In her model, design sits far upstream in product development as a strong business lever. It does not wait for a finished requirement and then produce screens. It helps define the solution in the first place. The experiment is whether that way of working brings business value, and whether the team can see it through.

She connects the shift to AI. As tools take on more of the production work, the role moves toward defining the human experience rather than designing screens. She expects designers to split in two directions: toward design engineering, where they build products directly, or toward deep specialization in craft, because in a world where everything looks the same, elevated craft becomes a real differentiator. Alongside that, she sees product and design roles merging, with design thinking about go-to-market and solution design from the start.

Small cohorts instead of classic A/B tests

For its big bets, Grubhub experiments by releasing features internally or to small cohorts and watching whether they win or lose. The goal is to act against business metrics without disrupting core revenue and order metrics while the answer comes in.

Michal's example comes from her post-purchase work. Grubhub saw a lot of customer care contacts, canceled orders and changed orders after checkout, driven by anxiety about whether the order was right. The team designed a review screen that appears right after an order is placed: is this the order you wanted, and is it going to the right place?

A screen like that could reduce care costs. It could also cost orders. So the team released it to a very small cohort and monitored both closely, confirming that care contacts fell and orders held before going further. Less A/B, as Michal put it, and more slow rollout to see whether an experiment wins.

The merchant concept test that changed direction

Asked for an experiment the team learned the most from, Michal described a recent one in Grubhub's merchant promotions space.

The team did the research first. They spoke with around fifteen merchants about how they use the product and what would help, and came away with a clear view of where merchants wanted to go: more metrics and more information.

Then they turned the concepts into something merchants could react to and put them in front of the people who would use them. Merchants preferred the much simpler experience.

"I'm glad we did it because I think we would have otherwise launched a very different product that probably wouldn't have won or wouldn't have been as successful," Michal said. The team pivoted and is now designing a different kind of dashboard.

Ashley recognized the pattern from his own product work. You spend months on something, it becomes your baby, and then a customer who is doing twenty other things at once shows you they have forgotten how it worked the last time they used it. Michal's response has become one of her working principles: product design is a master class in removing your ego from the experience. Users will always find ways to use a product differently than you expected, and the job is to build for what is best for the business and the user.

Prototypes in front of users, early and often

The concept test is not a one-off. Grubhub hosts a regular round table with merchants where the team brings products and features forward, gives merchants testing environments, and watches them work through a task. Can you get through this? Is it valuable?

AI has made that loop much faster. Michal's teams can throw prototypes together quickly and put real, clickable experiences in front of users early, running what she calls micro experiments throughout the process. The payoff is directional confidence before engineering commits, which is the whole point of de-risking upstream.

AI speeds up the other end of research too. Synthesis used to mean days of sticky notes. Now the team can run interview notes through NotebookLM or Claude and get a fast synthesis to work from.

Success metrics for every team in the room

When Grubhub does set up a formal experiment, Michal starts with success metrics, and she pushes teams to define them per discipline. How does product know it won? How does design know it won? How does engineering know it won?

For design, the metrics are about engagement and time on task, signals that the experience itself hit the mark. For the business and product side, the questions are whether the feature brings in more revenue, drives more orders and moves things forward for the business. Whatever the set is, it is agreed early and tracked through the experiment.

Why the numbers need a conversation

Ashley raised the case every experimenter knows. The research says there is a need, the feature ships, and engagement does not show up. Did users not see it? Did they see it and get confused? Was it simply not what they wanted?

Michal's answer is qualitative research after the experiment. Grubhub has the emails and user connections to reach people who went through a test, and the team incentivizes five or ten of them to talk through what happened and why they chose one option over another.

Even an unhelpful sounding answer is useful. A user who says they never noticed the button because they were dealing with a whiny kid and putting laundry away is telling you the feature is either arriving at the wrong point in the process or is not visible enough for someone half looking at their screen. That context is what turns a flat result into a variant that can win. Michal calls it the marriage of qual and quant, and she thinks too few teams build it into their post-experiment process.

Choosing where to de-risk

The conversation closed on a contrast. GrowthBook builds a technical product used across mobile, web, data warehouses and a long list of other tools, so Ashley's team ships a V1 quickly and innovates around the feedback, sometimes going from feedback in the morning to a fix in the afternoon. Grubhub, with millions of users, invests upstream so that it releases with confidence. A mistake there means a lot of hangry customers.

Michal does not think one approach is right. Both work, and the deciding factor is how fast engineering can move. If a team can release a product, see it underperform and change it within a sprint, iterating in production is a fine model. If not, the de-risking has to happen earlier.

What matters is choosing. At some point every team has to de-risk, either in a fast follow or upstream in the process. Her advice for any company setting up an experimentation program is to decide on that path intentionally and make it part of the product development process from end to end.

🎧 Listen to the full episode →

Table of Contents

Related articles

See All Articles
How MilliporeSigma lifted add to cart with a smarter search with Dorothy Crepin
The Edge Podcast
B2B product discovery: what MilliporeSigma learned rebuilding on-site search
PayPal's $180 million experimentation win with Gaurav Sethi
The Edge Podcast
PayPal's $180 million experimentation win
ServiceNow's Customer Zero approach to AI experimentation with Ashraf Karim
The Edge Podcast
ServiceNow's Customer Zero approach to AI experimentation

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.