Back to Podcast
A/B Testing
Culture
Testing AI

How Supercell A/B tests 300 million players without breaking trust

S1 | E36
Sep 1, 2026

On this episode of The Experimentation Edge, Ashley Stirrup talks with Shan Huang, data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, Brawl Stars, Hay Day, and Boom Beach. Shan explains how a famously decentralized, creative-first company with 300 million monthly active users runs fewer than 100 A/B tests a quarter and why that restraint is deliberate, how Supercell announces experiments to players in advance and promises make-up events to keep testing fair, and how importable AI skills now let anyone at the company analyze their own experiments, making quality consistency the next big challenge. It's for product managers, data scientists, and growth leaders balancing creative conviction with experimental rigor.

00:00 Intro
01:10 About Supercell and 300 million players
02:30 The central team and a decentralized culture
04:25 Fewer than 100 tests a quarter
05:40 Sharing learnings across independent game teams
07:35 Retention as the North Star
08:35 Onboarding experiments with gems and tutorials
12:45 Telling players about A/B tests
15:15 Hypotheses and proxy metrics
20:45 AI and the future of experiment analysis

"You can design a very creative and very good mechanism for players, but you don't really know what setup works best for the masses."

"If we make the experience of the game better, business outcomes come naturally. It's not our first goal."

"If you see the data contradicts with your hypothesis, that's actually the best part of the experiment, where you learn something completely new."

"Game players go to the game to have fun. If we are testing them by giving some people a disadvantage, it's not a good experience, because their main motive is just to have fun."

"Basically everyone can run and analyze experiments on their own now. But it also creates a new challenge: we don't know the quality of those AI-driven analyses."

The Experimentation Edge - Shan Huang
===

Shan Huang: [00:00:00] You can design a very creative and very good mechanism for players, but you don't really know what setup works best for the masses

Ashley Stirrup: Hello and welcome to today's episode. Today I'm excited to welcome Shan Huang, data scientist from Supercell. Great to have you, Shan

Shan Huang: Thanks for having me. Hello everybody

Ashley Stirrup: Why don't we kick things off by having you tell us a little bit about Supercell?

Shan Huang: Of course. Supercell is a mobile game company based in Helsinki, Finland. Supercell has developed a couple of very popular mobile games such as Clash of Clans, Clash Royale, Brawl Stars, Hay Day, Boom Beach, and we have a couple of more new games in the team

Ashley Stirrup: That's awesome. And you have an enormous number of users as well. Is that right?

Shan Huang: Yeah. So currently we have around three hundred million monthly active users. Historically, we've been also we have like mobile game in-industries like up and downs are [00:01:00] very dramatic. We have like even better times

Ashley Stirrup: Yeah. Yeah. That's still an incredible number. And how big is the company? Like, how many employees roughly?

Shan Huang: We've been expanding quite fast lately, the last two years or so. So right now we are about a thousand employees. Compared to other game companies or tech companies, still relatively small. But compared to Supercell before, we've been like really expanding a lot in the last two years.

Before that, it's I don't know, two years ago, maybe five hundred employees

Ashley Stirrup: Oh, that's a lot of growth then. That must be pretty exciting.

And can you tell us a little bit about your role there?

Shan Huang: So I'm a data scientist. Currently I'm in the central team. It's in the experimentation team. Our team's roles are two parts. One is building the experimentation platform so that the game teams can use the experimentation platform to run A/B test, run all kinds of experiments they [00:02:00] wanted to save the repeated effort of setting it up and analyze it and also making sure the analysis is rigorous.

The second part of our team's role is to drive the experimentation culture at Supercell. Maybe here I can talk about the unique culture of Supercell a bit, 'cause Supercell is unique in many ways. I guess the most unique way is Supercell is a very decentralized company, where decisions are made bottom up.

Each game teams are completely independent in what games they develop, what decisions they make in terms of their games. So historically A/B test was not part of the culture of developing a product because we consider ourself a very creative company, and we still are. So this creates a unique position of Supercell and also a unique challenge of how can we be more [00:03:00] hypothesis-driven while keep being creative.

Ashley Stirrup: Yeah. Yeah, that is a really interesting challenge. I, know a lot of teams, they feel like, a creative process is one that you should just be following that process. But even then, if you think you've done something amazing, you should test it to see if people are actually interacting with it and the impact of it.

So that, that's

gotta be quite a cultural change,

teaching people that way of thinking. Yeah. And roughly how many experiments are you running a quarter?

Shan Huang: At this stage, roughly... supercell have four or five big games and a few small games in testing. So this is our portfolio. And in total, we have less than a hundred AB test per quarter. So let's say ten to twenty per game, depends on how big the game is

Ashley Stirrup: Got it. Got it. Still, so that's still across six games, that adds up quite a bit

Shan Huang: Go

Ashley Stirrup: And can you tell us a little bit about experimentation at your company? Like how is it structured? How [00:04:00] much of it is centralized versus decentralized? Things like that.

Shan Huang: So we have a central team, which I'm part of the experimentation team who works across different game teams. We provide the tools, we help facilitate the culture. We also help as a partner to run those A/B tests correctly and analyze them. And then we have the game teams.

Each game team have their own data scientists, data analysts, product managers who decides... Ultimately, it's their game team's decision what to A/B test, how they want to structure the A/B test, and what decisions they want to make to roll out which variant and so on. Of course, central teams are a big part of the process, but ultimately it's the game team's decis-

Ashley Stirrup: Yeah. And with a decentralized model like that, I'd imagine the centralized team plays a pretty important role in making sure learnings get shared across teams?

Shan Huang: Indeed.

Ashley Stirrup: Yeah

Shan Huang: when game teams that's the price comes with [00:05:00] being independent and decentralized because you, as a game team member, you don't know what other games are doing. You are like laser-focused in your own game, and you have all the freedom to make the decisions within your own game.

And indeed it's a central team's challenge and role to make sure the learnings at Supercell are shared across different game teams

Ashley Stirrup: How do you try to facilitate that cross-team learning?

Shan Huang: We have a couple of attempts. So as a central team member, we proactively reach out to different game teams. So we have a relationship with all the game teams and their analysts and product managers. So we proactively talk to them, and if we see some insights that can be shared, we proactively share with them.

This is one channel. And we also have other more formal channels of sharing. For example, we've tried to have on a Slack channel where there is a bit more structured sharing per month that we ping different key stakeholders in each [00:06:00] game team and then they share a summarized version of their experimentation in the previous months, and then all the other people can see this thread.

And then in this Slack channel, people can als-also ask ad hoc questions about experiment setup or experiment learning interpretation they have on the fly, and other people can also see other people's reply. And then we have a insight library where we try to enforce a consistent documentation structure, and then game teams follow the same documentation structure put into same place, and then people can use this insight library to search through similar results.

With AI, this finding might be even easier

Ashley Stirrup: Yeah. Yeah, no, I think that's incredibly important 'cause it's so hard for humans to even read 10 experiments and keep it all in your head, let alone 100. So yeah. Can you talk a little bit about the business impact experimentation has had at Supercell?

Shan Huang: [00:07:00] Yeah. I think the most important thing at Supercell is player experience. So we believe business outcomes comes naturally if we can improve the player's experience. So we would say the most successful AB test would be those that can help us learn what makes our player engage more. So retention is our let's say North Star metrics most of the time, and we believe monetization revenue that is a part of the process.

If we make the experience of the game better business outcomes comes naturally. It's not our first goal.

Ashley Stirrup: That makes a lot of sense. You create a better game and the money will come as long as

you're creating happy users and growing your user base and things like that.

Shan Huang: That's right

Ashley Stirrup: So that leads right into is there an example of an experiment you ran where you had a lot of learnings?

Shan Huang: Yeah, I can share a few. let me start with there are many areas we can experiment at a Supercell portfolio levels. Of course, each game, they have [00:08:00] their stuff to experiment. And then we have the UA, which is a very data kind of numbers game. So there's a lot of experiment, a quasi-experiment done in the-- at the UA user acquisition channel.

And there's also... we also have a Supercell store, which is a website you can buy stuff, and then you have the common e-commerce A/B test there. But that's not the biggest focus. So maybe I can start with the experiment in games.

Ashley Stirrup: Sounds great

Shan Huang: yeah. I think maybe one maybe giving you a bit of context is that each game has a, has some onboarding process.

So we give you some tutorial, and we teach you the meta of the game, which is the system of the game step by step until you are familiar with all the mechanics of the game. One A/B test we've done in one of the game is that we we ask them to spend a large sum of gems right after we teaching [00:09:00] them that this is a premium currency.

The initial thought is that okay, we are teaching you this is... Gem is a premium currency. We teach you how to spend them. You get some better stuff back, and then later on, you might learn how to use those premium currency and give you a better experience later on. And what we are A/B testing is to remove this teaching step.

And it turns out that remove this step actually have a better retention, actually creates a better engagement for new players.

Ashley Stirrup: Got it

Shan Huang: the learning is probably that players are smart enough that to know that this is a premium currency and adding more teaching steps just creates friction without the benefit of letting them know what's what is it.

And we bel-- and the learning is players are smart enough to get this by themselves after they play with the game

Ashley Stirrup: Got it. So they learn pretty quickly that, "Hey, these points are important and I need to be careful how I [00:10:00] spend them," type of thing.

Shan Huang: Yeah. Yeah. Another one may be related is the tutorial steps. We've also done quite a bit of A/B test on the tutorial steps, like teaching like how do you teach each, let's say, battle step? How do you teach each action? What is the result of each action? T-through the A/B test, we could improve the completion rate of tutorial, for example.

But later, we learned that those those optimization didn't lead to a significant following day retention. Even though we improved the completion rate of tutorial, they don't lead to a better engagement on the next day. So the learning here is that we don't forget the big picture, which is ultimately, creating a better experience for the whole game loop, rather than just sub-optimally optimize like a few small steps which didn't lead to the big picture.

Ashley Stirrup: Got it. Yeah, that's an interesting one. My-- one of my [00:11:00] daughters worked for Activision for a summer, and so I tried the game that she was working on, and I was just lost. And so I could imagine how I would've appreciated tutorials. There probably were tutorials somewhere in the game that I just never found. But I guess part of what you're saying is make a better game, make it easier for people to understand how to do things in the

game and don't rely on the crutch of having tutorials. Is that kinda

what you're saying?

Shan Huang: that's the learning we had from those A/B test. But I guess also people by people, just population has diff-- We, we can only optimize for the average 'cause I guess we like different,

Ashley Stirrup: yeah. Yeah. I think over time, especially with AI, we have the potential to do a lot more personalization,

and I think that's gonna be one of the big, next unlocks in experimentation is how do you run experiments and really look at the segments and then create more personalized experiment-- features for those different segments.

I think there's a lot of opportunity there. But

Yeah,

Shan Huang: I agree. [00:12:00] Yeah.

Ashley Stirrup: Yeah. And I know that Supercell takes A/B testing very seriously and that your user community is very sensitive to it. Maybe you could talk a little bit about, like, how the, how you try to communicate what you're doing to the users

Shan Huang: That's right. We always try our best to be as transparent as we want to the player community. So if there are some A/B test coming up our community manager always try to ensure to tell players upfront that we are going to have testing or experimenting with some features in this area in the next month or two, because we are still not sure what design makes your experience the best.

So we want to try out a few things and learn from you, like what works best. But , we also want to make players rest assured that if they don't get the better experience during the test period, they will always have [00:13:00] a make-up event later on. So we take fairness very seriously.

Ashley Stirrup: It's such an interesting topic, because people hear testing and they think, "Oh, I don't wanna be tested on." And it's like we've got this new feature. 20% of the time these new features are better, and 80% of the time they're not. And so we sure would like to measure and know whether we're actually making your life better or not. But for the individual, if you don't understand all that, it can sound like a scary thing.

So sounds like you're investing a lot in that kind of communication and trying to be transparent and trustworthy with your user base.

Shan Huang: Yeah, game product is a bit different than other, commercial product, e-commerce. Because I worked in e-commerce before I joined Supercell, where users come to the website to buy something. But game players, they go to the game to have fun. They are trying to entertain themselves using this time.

So if we are testing them by giving some people a disadvantage, it's not a good experience [00:14:00] because their main motive is just to have fun.

Ashley Stirrup: Yeah, it makes total sense. So let's say you start working with a new product manager and they don't have a lot of experience with experimentation.

It's often a mindset change to th- realize that, gee, most of my experiments lose. And so if I'm running an experiment with the assumption that it's gonna lose, I wanna structure it so that I get the most learning possible from that experiment. And how do you kind of coach people

to think about experimentation through that lens and make sure they're designing it so they're gonna learn as much as possible?

Shan Huang: Yeah. I think the most important thing is to have the hypothesis written down before the test starts. Because if you write down the hypothesis in a clear structure, you form your hypothesis better. Let's say if I changed this feature, like by this, I would assume this metric go up or go down by this percent because I did some preliminary study or [00:15:00] because my previous experience in the, let's say the last year's experience .

So you write down your hypothesis, and then you run your test, and then you collect your data. If you see the data contradicts with your hypothesis, that's actually , the best part of experiment where you actually learn something completely new because it contradicts with your previous experience.

And then there you can look at more detailed breakdown metrics or look at more segments to understand further where does that hypothesis contradicts with data. You can also look at qualitative data like you can give user surveys or have player interviews to really watch and ask them why they behave or feel this way so that you learn and iterate your product

Ashley Stirrup: Yeah. Yeah. And so you mentioned that retention's kind of your North Star metric, but how do you think about key metrics in general?

Shan Huang: So each [00:16:00] game have a North Star metric, as you said. It's usually retention. Sometimes it's also LTV, lifetime value, like however it's defined. But those are really long-term metrics. We cannot afford to run experiment for a year or longer. Usually, the experiment only lasts for a week or two, like a month at most.

So what we want to do is to find those proxy metrics that is sensitive to change by the feature during the experiment, and also has a relation to the North Star metric. So if the North Star metric is long-term retention, let's say we assume that daily playtime is a proxy to that. The first three days of daily playtime is a proxy to long-term retention.

And we also know by changing the gameplay, it will have an effect on the first three days of daily playtime. That's the ideal case, but of course, reality is [00:17:00] difficult. It's like you can't always find such good metrics that is both sensitive to change and are proven to have a causal relationship to the North Star metric.

So in reality, it's always a iterative process with uncertainty. You have a KPI tree. You divide the North Star into like many layers and detailed metrics, and then you iterate your product learning from various areas to learn, like what changes could be sensitive to what metrics, and do they have a real relationship to the North Star?

It's always an ongoing process.

Ashley Stirrup: Got it. And have you built out models that predict long-term, lifetime value and like the effect of an experiment over time?

Shan Huang: At Supercell, we don't have this at the central team, at least yet. I don't know if the game teams might have some ad hoc analysis. At the central team, we haven't built this projection into the product yet. Before Supercell when I work at e-commerce [00:18:00] company, we did have a attempt to project long-term impact based on the short-term metrics in the experimentation period.

Basically, we take some features during the experiment and use a machine learning model to predict how does the impact last after months using trained by historical data. I would say take the result with a grain of salt, at least from my experience at that time, four or five years ago.

Ashley Stirrup: Yeah.

Yeah

Shan Huang: inaccurate and, yeah, the product is changing as well, so it's hard to really have a very accurate long-term prediction

Ashley Stirrup: Yeah. Yeah. It's so interesting how things can perform well in the moment and then, why do they maybe not show that same results over time? And yeah I think that's an area there's a lot of opportunity for future learning for the whole industry as to how we get better at that. And that actually leads into our last question here is, how do you see experimentation evolving at Supercell?

Shan Huang: I think [00:19:00] over the last two years, the mindset at the company has been evolving. Previously, we are very decentralized, and we have, I think, we take pride in being a completely creative-driven industry and game company. But I think over the last two years or so, or maybe even more, we have been evolving the mindset a bit as the company has been bigger and also the player base is really large.

You can design a very creative and very good mechanism for players, but you don't really know what setup works best for the masses. Like what is the-- works best for the average of millions or billions of people. I think there are ways that art and creation works, and there are also areas where experimentation and being rigorous, being hypothesis-driven work better.

So w-we are, like, evolving our culture, and we are having a momentum of finding what are the right things to experiment that [00:20:00] helps it all.

Ashley Stirrup: Yeah, it makes a lot of sense. It sounds like culturally you've already come a long way and that you're continuing down that path of

Baking in a data-driven approach on top of your kind of core creative approach.

Is that a good way to say it? Yeah, yeah.

Shan Huang: I think we prefer to use the word hypothesis-driven

Ashley Stirrup: Yeah. Yeah

Yeah. Absolutely. What about AI? Do-- how do you see AI coming into your experimentation approach?

Shan Huang: It's a really new era we felt 'cause now everybody can use AI to run their own analysis as long as they set up the correct way and they can just give a... We have some experiment skills they can just import to Claude, and then they can ask Claude, "Okay, use this skill and now I have been set up this experiment and data.

You can find the data yourself and analyze it. Tell me the result. What shall I do?" The good part is it's like more and more people are [00:21:00] easy to run A/B tests and analyze by themselves. It's not restrict to only data analyst or product manager with experiment knowledge. Basically everyone can run and analyze experiment on their own now.

But then it also creates a new challenge that we don't know the quality of running those AI-driven analysis. I'm sure that people with experience, they can guide the AI to run the analysis correctly, but we don't know everybody who's running analysis using AI.

So that creates a new problem. How can we make sure the quality of AI analysis is consistent across all the usage?

Ashley Stirrup: Yeah. I think that's such a great topic 'cause like rigor is just so unbelievably important in A/B testing. Obviously you don't wanna fool yourself into thinking something that's a winner that's actually a loser or vice versa and then maybe you stop investing in an area you should have been investing in. But [00:22:00] also, obviously at some point it can undermine trust. And the challenge with AI is that it can give you a very confident answer, but with only half the data. And so making sure that you're applying AI in the right ways is

incredibly important to unlocking a stronger self-service experience.

That's something we're investing a lot in at GrowthBook, how do you create like experiment templates and making sure you've got the right guardrails and, doing the right experiment design. Leveraging AI for those things, but doing it with a lot of guardrails to make sure it's like your best data scientists would do.

So that to me, I think is just a super exciting opportunity for the market to invest in. So Shan, thank you so much for being on the show today. Obviously gaming is fun by itself, but it's especially fun talking about A/B testing for gaming, so really enjoyed having you on.

You brought up a lot of great topics today.

Shan Huang: Same here. Thank you for having me, Ashley.

Thank you .

[00:23:00]

About Shan Huang

Shan Huang is a data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, and Brawl Stars. He builds the experimentation platform and champions hypothesis-driven culture across Supercell's independent game teams, bringing earlier e-commerce experimentation experience to testing at 300-million-player scale.

LinkedIn
Shan Huang
Role
Data Scientist
Industry
Media & Gaming

Subscribe to the podcast

Takeaways from this conversation

AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company.

S1 | E36

Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged.

S1 | E36

Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner.

S1 | E36

In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams.

S1 | E36

Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume.

S1 | E36

Top takeaways from other favorite conversations

All Takeaways

Losing tests often create more value than winners because they stop expensive mistakes before they ship.

Go to S1 | E14
Theme
ROI
Role
Product
Industry
Financial Services
Featured
false

Executive engagement is real at Home Depot: leaders join 30-minute readouts, search the experiment library, and ping analysts directly because they treat A/B testing as the golden rule for measuring incrementality.

Go to S1 | E21
Theme
Culture
Role
Data Scientist
Industry
Retail
Featured
false

Measure DORA metrics and developer sentiment; remove mundane toil to increase speed and satisfaction.

Go to S1 | E5
Theme
Velocity
Role
Exec
Industry
Financial Services
Featured
false

Faster is not always better. Fin raised latency artificially and positive feedback went up, likely because a small delay makes an AI feel like real work.

Go to S1 | E28
Theme
A/B Testing
Role
Data Scientist
Industry
Business Tech
Featured
false

A feature that fails early in a flow can succeed later; placement and timing often matter more than the idea itself.

Go to S1 | E14
Theme
A/B Testing
Role
Product
Industry
Financial Services
Featured
true

When senior leaders push ideas, Massey's team tests them instead of arguing—then delivers results that either validate the idea or identify three better alternatives the data actually supports.

Go to S1 | E11
Theme
Culture
Role
Exec
Industry
Logistics
Featured
false

Experimentation short-circuits political debates by removing opinion from product decisions.

Go to S1 | E13
Theme
Culture
Role
Exec
Industry
Business Tech
Featured
false

Build an AI ecosystem with clear purposes (productivity, engineering, consumer) and a steering committee to avoid duplication.

Go to S1 | E5
Theme
Culture
Role
Exec
Industry
Financial Services
Featured
false

Start with low-risk, high-yield AI use cases—unit tests, documentation, and security triage—to build confidence and momentum.

Go to S1 | E5
Theme
AI-Native Dev
Role
Exec
Industry
Financial Services
Featured
false
The experimentation edge podcast logo with a picture of host Ashley Stirrup