The Edge Podcast

How Zalando connects every experiment to its North Star

How Zalando connects every experiment to its North Star

Guest: Mi Tian, Head of Applied Science, Zalando. Show: The Experimentation Edge with Ashley Stirrup. Company: Zalando, the Berlin-based e-commerce platform operating in around 30 European markets with over 60 million monthly active users.

Zalando runs more than 1,000 experiments a year. User randomization tests, geo experiments, marketing experiments, personalization, recommendations, size and fit. The whole customer funnel is under continuous optimization, across many teams with different needs, velocities, and challenges.

At that scale, the hard problem isn't running experiments. It's making sure every one of them is worth running. On The Experimentation Edge, Mi Tian, Head of Applied Science at Zalando, explained how her team keeps 1,000 experiments pointed at the same destination: the company's North Star goal.

🎧 Listen to the full episode →

An experimentation org that reports to the CFO

Mi's team sits in an unusual place. She leads applied science within Zalando's economics and experimentation org, a central decision-making group directly under the chief economist and CFO.

The setup is deliberate. The org combines economists with experimentation-focused data scientists, and the shared mandate is decision-making, not any single method. "We try to use the right approach to address fitting questions," Mi explained. When online experiments are an option, they run them. When they aren't, the team leans on causal inference and other economics approaches to measure incremental impact.

That executive positioning matters culturally too. Zalando's CFO is an active champion of experimentation, regularly asking the team what she can do to better support the effort. As Ashley noted, when someone senior is paying attention to results and asking questions, it signals to the whole organization that this work matters, and buy-in follows.

The KPI tree: working backwards from the North Star

The core of Zalando's design discipline is a simple rule: any product goal needs to contribute to Zalando's North Star goal.

The obvious problem is that a North Star metric is nearly impossible to move, or even read, inside a single A/B test. Zalando's answer is a process Mi calls growing a KPI tree, or cascading outputs to controllable inputs. It works backwards in three steps.

First, start with clear goals and hypotheses. What product change is the team creating, and how is it expected to impact users or the business? Is it meant to improve short-term transactions? Drive engagement? A good experiment starts here, not in a testing tool.

Second, map the goal down to a primary success metric. The team selects proxy metrics that are sensitive enough to be effectively measured in the team's own A/B tests, while still connecting upward to the North Star. This is the critical middle layer. Without it, teams either chase metrics that are measurable but meaningless, or meaningful but immovable.

Third, map the metric down to controllable inputs. These are the day-to-day product changes the team is actually working on. If the chain holds, a shipped change is expected to move the proxy metric, and the proxy metric is expected to move the North Star.

From there, the team defines supporting observational metrics to help interpret results, sets standards like minimum detectable effect, and assesses feasibility before the experiment enters the roadmap or backlog. Goals first, then metrics, then configuration. On the platform side, Zalando's engineering team built a self-service experimentation platform, so any team can launch an experiment and analyze results end to end through a self-service dashboard.

Ashley summarized the practical payoff neatly: identify the first metric you'd expect a new feature to change. If you don't move that first metric, you're never going to move the North Star. The KPI tree turns that intuition into a repeatable design step.

A portfolio, not a queue

Design rigor alone can push a program toward timidity. If every experiment must justify itself against the North Star, teams may retreat into safe, incremental tests. Zalando counters this with deliberate portfolio balance.

Some experiments help the business stay in the game: standard confirmatory tests aimed at improving existing features, where confidence is high. Others are game-changing bets. "Those are riskier," Mi said. "They could give us negative results, but it could also change the game and help us win big."

The mechanism that makes room for both is risk tiering. When defining the experiment roadmap, the team assesses each experiment's risk tier and assigns a matching measurement strategy. High-risk experiments, such as disruptive feature designs that could move financial metrics, or long-duration holdouts that burden both users and the engineering stack, get heavier rigor. Lower-risk experiments move fast. The goal, as Mi put it, is to make rigorous decisions without slowing down the decision-making process too much.

The discovery feeds: a game changer, measured properly

The clearest test of this philosophy came with one of Zalando's biggest recent launches. Last year, the company introduced discovery feeds, a major reinvention of the homepage. The traditional experience, sorted by categories like men's, women's, and kids' fashion, evolved into a rich, entertaining experience with personalized feeds, video streams, boards, and curated content.

Here's the striking admission: at launch, the team didn't know how to measure success. They knew they wanted to scale the product, drive engagement, and drive conversion once it matured. But an experience this new had no established success metric.

Rather than guessing at one, the data scientists and product teams designed a comprehensive measurement plan that ran from day one of the launch. Short-term A/B tests measured the impact of small product changes quickly, feeding the iteration cycle. Long-term holdouts measured the incremental impact of the feeds product as a whole, from the first day.

After a few months of iteration, the results came in: strong engagement metrics, no damage to short-term transaction metrics, and positive impact on gross merchandise value. A genuinely risky bet, and the measurement system was strong enough to prove it paid off.

Where it goes next: agents, simulation, and a changing shopper

Experimentation at Zalando keeps evolving because the product space keeps evolving. Customers arrive with expectations shaped by TikTok-style experiences. Curated boards blend B2B and B2C into a single surface, with brands and fashion influencers publishing content inside the shopping experience. Each new surface raises the same question: how are you going to test for it?

One frontier Mi's team has started exploring is LLM-based agents that simulate real customer behavior. She is careful about the claim: "We're far from replacing real A/B tests with agents." The nearer-term opportunity is a cheaper, faster simulation layer for screening hypotheses, so teams can prioritize the high-impact ones before spending real traffic. Simulation could also help where live testing isn't feasible at all, such as Cyber Week, Black Friday, or specific locations and product types that can't carry experiments.

The hypothesis space is bigger than ever, and for Mi, that makes culture the real flywheel: continuously feeding experiment insights back into the product development cycle.

She closed with an analogy that stuck. She misses the old Saturday afternoons at the mall: checking familiar brand shops, wandering into unfamiliar ones, trying on jeans, stopping for dinner, catching a movie. E-commerce platforms, she argued, are now responsible for making shopping that engaging, convenient, and entertaining. "It is a longstanding business, but we have new missions, we have new challenges."

For a mature industry, that's the point worth remembering. The checkout flow being 1% better was never the ceiling. The teams that will win are the ones with an experimentation practice strong enough to measure the game changers, and a KPI tree that keeps every bet, safe or bold, connected to the goal that matters.

🎧 Listen to the full episode →

Table of Contents

Related articles

See All Articles
Realtor.com on using your AI as a junior data scientist with Whitney Perez
The Edge Podcast
Realtor.com on guardrail metrics, the 30/30/30 rule, and AI as your junior data scientist
Learneo on testing the opposite of every hypothesis with Rich Liebling
The Edge Podcast
Learneo on testing the opposite of every hypothesis
The Edge Podcast
The four questions Early Warning asks before any A/B test

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics — free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.