Back to Podcast
A/B Testing
Scale
Future of Testing

How Zalando connects every experiment to its North Star

S1 | E39
Sep 10, 2026

Summary

How do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition.

📝 Read the full blog post →

Chapters

00:45 About Zalando and its marketplace model
02:00 Mi's path from engineering to experimentation
03:10 Economists and data scientists in one decision-making org
04:15 Running over 1,000 experiments a year
05:55 What makes an experiment high risk
07:10 Sharing learnings through standardization and champions
09:10 The discovery feeds launch and its measurement plan
11:15 Balancing the experimentation portfolio
12:35 Growing a KPI tree from the North Star
16:05 LLM agents and the future of experimentation at Zalando
18:35 New missions for a longstanding business

Notable Quotes

"A good experiment starts from clear goals and clear hypotheses. What product change do you want to create, and how is it expected to impact our end users or our business?"

"We try to really balance the experimentation portfolio to not limit ourselves on the small experiments only."

"Some others are game-changing ones. Those are riskier. They could give us negative results, but it could also change the game and help us win big."

"We're far from replacing real A/B tests with agents. But maybe we can get a cheaper and faster simulation layer just to screen all these hypotheses so we can prioritize the high-impact hypotheses better and run them on the real traffic."

"It is a longstanding business, but we have new missions, we have new challenges."

Transcript

The Experimentation Edge - Mi Tian ===

Mi Tian: [00:00:00] I think a good experiment starts from good goals, clear goals and clear hypotheses. You need to have this idea in your mind. What product change you want to create this time, and how is it expected to impact our end users or our business? Is it going to, improve short-term transaction?

Is it going to drive engagement? You need to have this high-level goal in mind, and then we tend to help them to map this goal to a company North Star goal.

Ashley Stirrup: Hello, and welcome to today's episode. I'm excited to have Mi Tian, head of applied science, economics, and experimentation at Zalando. Welcome to the show, Mi.

Mi Tian: Thank you so much, Ashley. And hey, everyone. Really happy to be here.

Ashley Stirrup: Yeah. Excited to have you on the show. It's Zalando's an amazing business. Maybe we could start with you just telling us a little bit about the company.

Mi Tian: Sure. Zalando is a German e-commerce company headquartered in [00:01:00] Berlin, where I'm based. A really beautiful and vibrant city. It specializes in beauty, fashion, footwear, all kinds of assortment. Running B2C and B2B businesses serving as a platform for merchants and partners to sell their products and connect with our customers.

I believe it operates across Europe in around 30 European markets and countries with over 60 million monthly active users.

Ashley Stirrup: Yeah, that's pretty impressive. And then it's a bit of a hybrid model where you're both selling directly and then you have marketplace sellers as well, right?

Mi Tian: Creating a lot of interesting challenges for experimentation enthusiasts.

Ashley Stirrup: Yes, no question about it. Marketplaces definitely. Now you suddenly are optimizing for three different groups, right? You're optimizing for the end customer, for the sellers, and then for your own business as well. Terrific. Why don't we start with a little bit on your background? You've been doing experimentation for some [00:02:00] time now.

Mi Tian: For quite some years. My background, my education is actually in electronic engineering. But then after my PhD, I got more and more into data science, understanding patterns from data. So I started working in data science, first as, like a research engineer type of data scientist, and later on focusing more on user engagement, user satisfaction, and how do we measure the impact a lot of the times with experimentation.

Right now I'm leading this applied science team in the economics and experimentation org in Zalando. It's a very central decision-making org directly under our chief economist and CFO. So we're at a really great position to to observe and harness a lot of data and to really understand how might we make data-driven decisions and help [00:03:00] the company, help the teams to steer success to steer ROI.

Ashley Stirrup: Yeah. It's interesting that you have economist in your team or org. Often these groups are really pretty focused on experimentation, maybe user experience. So how come the economics piece?

Mi Tian: So yeah, that, that's a really interesting setup in our org. we're we are part of-- we have economists and more experimentation-focused data scientists. The goal for us is really decision-making. We try to use the right approach to address fitting questions. Sometimes when experimentation is an option, then we of- we often opt for running online experiments.

Sometimes when not, we try to use leverage causal inference or other economics approaches to generate in insights to measure incremental impact. So we try to use different methods, different approaches to really solve the [00:04:00] questions and make decisions.

Ashley Stirrup: Got it. Super interesting. And can you tell me a little more about experimentation at Zalando? How big is the team? How much is it centralized versus decentralized? Things like that.

Mi Tian: That's a great question. Many teams run experiments at Zalando. At the moment, it's like we run probably around 1,000, over 1,000 experiments per year. These are really different types of experiments. Some are user randomization, some are geo experiments, some are targeted for marketing improvement, some are, like, more user-focused, personalization, recommendation, size and fit.

It's a whole conversion funnel, customer funnel that we're trying to optimize for, and then we run different types of experiments across all teams. Different teams have different needs and [00:05:00] velocity and challenges. So in our team, we try to really hold ourselves accountable and understand the the problems from these teams running experiments and try to provide value.

We often work on methodology, define templates, work on standardization in order to drive the measurement strategy and measurement practice. Sometimes we also embed as data scientists in those teams ourselves to execute some high-risk, high-stake experiments ourselves, working very closely with product teams.

Ashley Stirrup: Got it. What makes you consider some experiments higher risk than others?

Mi Tian: So sometimes we often use financial metrics as the key business KPIs in many of our experiments. So having very disruptive feature design that impact user experience a lot or [00:06:00] even have potential impact on our financial metrics, these are all considered high-risk experiments..

Some experiments are designed to be, like, long-duration, long-term holdouts. This could pose risks as well, not only on the user end, but also on the engineering stack because you need to hold certain version of the code that the experience of your product up and running. So yeah, there are different types of high-risk experiments, and we try to when defining experiment roadmap and portfolio, we try to assess their risk tiers to come up with different measurement strategy and risk measurement strategy to be able to make rigorous decisions but not like slowing down the decision-making process too much.

Ashley Stirrup: Yeah. Yeah, that makes a lot of sense. How do you, since you're supporting so many different teams, how do you try to help kind of share [00:07:00] learnings across the organization?

Mi Tian: I think we work in a hybrid model, and that really helps. Hybrid model as we focus both on the foundational methodology piece, but also work hands-on in leading some of those experiments. In the meantime, we work with those, we call them experimentation champion teams or champion groups, who are data scientists or product teams that are really in experimentation.

We try to leverage their impact to amplify the good practice as well. Firstly, we try to work on standardization. We define measurement strategy, incremental impact measurement. We standardize what metrics we track and measure and report, what common language we speak when communicating impact.

So when we report impact in this shared [00:08:00] language, teams tend to really pay attention to the most important metrics and goals. So this is one thing, standardization. Second is hold ourselves accountable to lead and run experiments ourselves, and really just to set, try to set the good practice.

And last but not least, just to work closely with product teams and let them leverage their great impact, their product impact to amplify the value of running effective experiments.

Ashley Stirrup: Yeah, and I think you said that the CFO is a, a pretty big champion of experimentation in-inside the company too.

Mi Tian: She's really awesome.

Ashley Stirrup: That's great. Yeah, I find it makes such a huge difference when you have someone senior in the organization who's really paying attention to the results and asking questions, trying to learn.

It shows everyone else how important it is, and that just helps create buy-in across the org.

Mi Tian: Yeah. And she often actively asks us [00:09:00] questions like, "What can I do to support you better to drive this effort, this initiative?" So really appreciate that.

Ashley Stirrup: Oh, that's great. That's great. Can you give us an example of an experiment you've run where you had a lot of learnings, maybe something that was a loser that had some surprising results?

Mi Tian: That's a great question. I think, I think in many learning experiments or launch experiments, we harvest a lot of learnings. One example could be probably the last year Zalando launched the discovery feeds. It's a really big change on the homepage. The traditional sort like male, female, kids fashion, the traditional experience was evolved into this rich experience with personalized feeds, video streams, boards, and curated content.

So it's a very entertaining and [00:10:00] engaging experience. And for us, we didn't really know how to measure success in the very beginning, apart from knowing that we want to really scale this product, drive user engagement, and drive conversion when it's mature. The data scientists, together with experienced product teams designed a very comprehensive measurement plan with short-term AB tests and long-term holdouts from the day one of the launch so that we can measure the very small product changes in a fast manner, but also measure the incremental impact of the feeds product from day one.

So there was a lot of interesting learning for us.

Ashley Stirrup: Interesting. I would imagine with something like that, we talked about risk earlier. A lot of risk on that one that could be a fabulous new experience for users that grows the business, or it could actually hurt the business.

Mi Tian: It could. When we set up experimentation portfolio, we [00:11:00] also balance different types of experiments. Some are helping us to stay in the business. Those are more standard confirmatory experiments which are aimed to improve existing features.

Those we tend to have higher confidence in. Some others are game-changing ones. Those are riskier. As you said, they could give us negative metrics negative results, but it could also change the game and help us win big. So we try to really balance the experimentation portfolio to not limit ourselves on the small experiments only.

Ashley Stirrup: Yeah. Yeah, I think that's a great example of a business being willing to take risks, so that's pretty exciting. So how did it do overall? Was it a successful experiment?

Mi Tian: I think the overall measurement strategy was super successful. After a few months We had really strong engagement metric result, but also we had, when we didn't hurt short-term transaction metrics. And we actually [00:12:00] had positive impact on gross merchandise value and type of metrics.

So it was exciting

Ashley Stirrup: Oh, wow.

Mi Tian: from the whole team

Ashley Stirrup: That's terrific. That's great. And so let's say you're working with somebody who's relatively new to experimentation. How do you help them think about crafting an experiment so they get the most learnings? Like, how do you make sure they design it and they're thinking about, "Oh, if this loses, what are my next questions gonna be?"

Things like that.

Mi Tian: I think a good experiment goals, clear goals and clear hypotheses. You need to have this idea in your mind. What product change you want to create this time, and how is it expected to impact our end users or our business? Is it going to, improve short-term transaction?

Is it going to drive engagement? Or you need to have this high-level goal in mind, and then we tend to help them to map this goal to a company North Star goal. [00:13:00] Any product goal needs to contribute to Zalando's North Star goal. But how? We then help them to work backwards from this top-level goal to map this goal to controllable inputs.

Sometimes we call them growing a KPI tree or cascading the outputs to controllable inputs. So in this process, we start from Zalando's North Star goal. We try to map down to a primary success goal for that product by selecting the more effective or sensitive proxy metrics that they can effectively measure in their own A/B tests.

And then we try to then map this primary success metric down to controllable inputs. These controllable inputs are the day-to-day work, the product change that the team is actually working on. So we ensure that the product change that we're going to run the experiment for is [00:14:00] expected to impact Zalando's North Star goal.

In how, then we then define additional observational metrics, support metrics, in order to help them interpret the results in the end, and also the standards likeMDE or other configurations. So we try to have this playbook defined or the template defined for them to first focus on the goals, and then we define the metrics, and then we configure the experiments based on the selected metrics and understand how feasible this experiment is before putting that into our experimentation roadmap or backlog.

So this is more from an experimentation design end. From a platform end, our super awesome engineering team also built this very self-service experimentation platform. It's pretty straightforward end-to-end experience for one to launch an experiment and analyze the results using the [00:15:00] self-service dashboard.

Ashley Stirrup: Got it. Got it. Yeah, I think it's always interesting to kinda start with your North Star metric, but then also look at the new feature you're introducing and say, "what's the first metric I would expect to change?" Will I see more clicks? Will I see more engagement? Things like that.

And then, if you don't move that first metric, you're never gonna move the North Star metric. But then, kinda create a chain that allows you to see, okay, I got this, then I got that, and okay, now I got revenue. I think that's always super interesting to make sure you're thinking about all that, and then thinking about which population.

Like, maybe my power users will respond better to this feature than my new users, that type of stuff, and understanding at that sub-segment level how things perform. Maybe we could wrap up with you talking a little bit about how you see experimentation evolving at Zalando.

Mi Tian: So I've been at Zalando for almost two years, and I see the culture and the practice evolving every day. It's like the [00:16:00] velocity is growing so much and the product space is changing much. Right now in this GenAI era, we have... customers have different expectations for the product, for the platform, and we also have different tools.

One recent advance we started looking more into is leveraging LLM-based agents. For example, what if we use these agents to simulate real customer behavior? We're far from replacing real AB tests with some agents. But maybe we can get cheaper and faster simulation layer just to screen all these hypotheses so we can prioritize on those high-impact hypotheses better and run them on the real traffic.

Or we can use these approaches when live testing is not feasible. Cyber Week, Black Friday, are certain locations or certain types of products that you just can't run experiments on. [00:17:00] And the product experience is changing so, so much. The business model is also evolving.

For example, now you have all these fashion trends and influencers and people are... Like, users are getting accustomed to, the TikTok or these kinds of experience, and then that they, that changes their expectation or your behavioral patterns on a e-commerce platform as well.

Take the boards, for example. Curated boards, you have brands, you have curated content, fashion influencers in there. You can say it's a B2B experience integrated into a B2C experience. So all these things, how are you gonna test for them? So I think the hypotheses, the problem space itself is being... It's so big now. It's just getting really exciting, and that comes to the culture part. We always need to continue to evolve the experimentation culture and then really spin the flywheel to just to always feed the [00:18:00] insights to the product development cycle. So the culture is all exciting.

Ashley Stirrup: So interesting, 'cause you think of e-commerce as a pretty mature industry at this point and yet, like the TikTok example's a fabulous one, where things continue to evolve and new innovations keep coming along. And so it's not just how can I make this checkout experience 1% better, but it's great.

Yeah,

Mi Tian: I always miss those good old days when you just spend a whole Saturday afternoon in a shopping mall, right? You check out those familiar brands shops. You visit a few unfamiliar ones aimlessly. You fit some jeans on, or you stop by the restaurant section and have dinner, and then in the end you're in the mood to watch a movie in the mall.

I feel nowadays the the e-commerce platforms are responsible to like really make this shopping experience engaging and easy and convenient and fun, [00:19:00] entertaining. So for me, it's e-commerce. It is a longstanding business, but we have new missions, we have new challenges.

Ashley Stirrup: Yeah, I love that. Those are really great examples. I like that kind of comparison to the mall experience where it's not just shopping, but it's a whole experience. So

Mi Tian: Discover, you entertain yourself, yeah. And you really connect yourself to the brands, to your lifestyle that you like.

Ashley Stirrup: Terrific. Mi, it was so great having you on the show. It's just so clear that Zalando's doing a number of amazing things in experimentation and we just really appreciate you sharing all of your experiences.

Mi Tian: Thank you so much. I learn a lot from your podcast.

Ashley Stirrup: Oh, thank you . You made my day.

[00:20:00]

About Mi Tian

Mi Tian is Head of Applied Science at Zalando, part of the central economics and experimentation org reporting to the chief economist and CFO. With a PhD in electronic engineering and a background in data science and user engagement, she helps steer more than 1,000 experiments a year, connecting each one to Zalando's North Star through KPI trees and risk-tiered measurement strategies.

LinkedIn
Mi Tian
Role
Data Scientist
Industry
Retail

Subscribe to the podcast

Takeaways from this conversation

No items found.

Top takeaways from other favorite conversations

All Takeaways

JavaScript injection tools carry hidden costs: broken pages, inconsistent results, and rework to reclaim your own data for deep dives.

Go to S1 | E30
Theme
A/B Testing
Role
Product
Industry
Marketplace
Featured
false

Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top.

Go to S1 | E23
Theme
ROI
Role
Engineer
Industry
Marketplace
Featured
false

Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test.

Go to S1 | E21
Theme
A/B Testing
Role
Data Scientist
Industry
Retail
Featured
false

DART measures behavior, not opinions. Four signals read off logs and transcripts: decay, acceptance, relevance, and task completion.

Go to S1 | E15
Theme
Testing AI
Role
Product
Industry
Business Tech
Featured
false

Treat ML features as living systems: feature-flag rollouts, realistic staging, drift monitoring, and LLM-as-judge evaluations—and be willing to kill “wins” that erode trust.

Go to S1 | E3
Theme
Feature Flags
Role
Exec
Industry
Marketplace
Featured
false

A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control.

Go to S1 | E34
Theme
A/B Testing
Role
Product
Industry
Consumer Tech
Featured
false

Plan for failure before you run a test. A pre-built playbook for a loss prevents confirmation bias and keeps teams from gaming the metrics.

Go to S1 | E19
Theme
Culture
Role
Exec
Industry
Financial Services
Featured
false

Run a broad explore experiment first; small, over-narrowed populations lack power and raise the odds of a false negative. Find the responsive segment with heterogeneous treatment effects afterward.

Go to S1 | E18
Theme
A/B Testing
Role
Data Scientist
Industry
Media & Gaming
Featured
false

Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis.

Go to S1 | E37
Theme
A/B Testing
Role
Exec
Industry
Financial Services
Featured
false
The experimentation edge podcast logo with a picture of host Ashley Stirrup