A/B testing 300 million players without breaking their trust: the Supercell approach

Three hundred million people play a Supercell game every month. Clash of Clans, Clash Royale, Brawl Stars, Hay Day, Boom Beach: some of the most successful mobile games ever made, built by a company of about a thousand people in Helsinki. A player base that size could justify thousands of A/B tests a year, and plenty of companies a fraction of Supercell's scale run exactly that many.
Supercell runs fewer than a hundred a quarter.
That number is not a program falling behind. It is the most deliberate thing about how the company experiments, and understanding why requires understanding two things Supercell refuses to compromise: its creative culture and its players' trust. On The Experimentation Edge, Shan Huang, a data scientist on Supercell's central experimentation team, walked through how both survive contact with rigorous testing, and what happens now that AI lets anyone in the company analyze an experiment on their own.
A company where decisions are made bottom up
Supercell's structure explains almost everything about its testing philosophy. "Supercell is a very decentralized company, where decisions are made bottom up," Shan explained. "Each game teams are completely independent in what games they develop, what decisions they make in terms of their games."
Shan's central experimentation team has two jobs. The first is the platform: giving game teams a shared way to run and analyze A/B tests without repeating setup work, with rigor built in. The second is harder to put on a roadmap. "The second part of our team's role is to drive the experimentation culture at Supercell," he said. Because for most of the company's history, A/B testing simply wasn't how games got made. "Historically A/B test was not part of the culture of developing a product because we consider ourself a very creative company, and we still are."
That framing produces the question that anchors the whole episode: "How can we be more hypothesis-driven while keep being creative."
Note the word choice. Not data-driven. When Ashley described the company as baking a data-driven approach on top of its creative core, Shan offered a gentle correction: "I think we prefer to use the word hypothesis-driven." The difference is not semantic. Data-driven can mean chasing whatever the dashboard rewards. Hypothesis-driven means a designer's creative conviction comes first, gets written down as a falsifiable claim, and then meets the players. In Shan's telling, the contradiction is the prize: "If you see the data contradicts with your hypothesis, that's actually the best part of experiment where you actually learn something completely new."
So the volume stays low on purpose. Four or five big games plus a few in testing, ten to twenty experiments per game per quarter, each one chosen because the answer matters. Retention is the North Star, and revenue is deliberately not the goal: "If we make the experience of the game better, business outcomes comes naturally. It's not our first goal."
The scale is what makes the discipline necessary rather than optional. "You can design a very creative and very good mechanism for players," Shan said, "but you don't really know what setup works best for the masses." Intuition produces the idea. Only an experiment can tell you how it lands across the average of hundreds of millions of people.
Telling players they're in the test
Here is where Supercell diverges from nearly every experimentation program you've heard of: they tell players about A/B tests before running them.
"We always try our best to be as transparent as we want to the player community," Shan said. When a test is coming, community managers announce it upfront: the team will be experimenting with features in this part of the game over the next month or two, because they are not yet sure which design makes the experience best. The message to players is an invitation, not a disclosure. We want to try a few things and learn from you.
Then comes the part that turns transparency into policy. "We also want to make players rest assured that if they don't get the better experience during the test period, they will always have a make-up event later on. We take fairness very seriously."
Why go this far? Shan's answer draws on his previous life in e-commerce. "Users come to the website to buy something. But game players, they go to the game to have fun. They are trying to entertain themselves using this time. So if we are testing them by giving some people a disadvantage, it's not a good experience because their main motive is just to have fun."
A shopper who lands in a losing checkout variant loses a few seconds. A player who lands in a losing game variant loses some part of the thing they came for. Supercell's community is famously vocal, and "I don't want to be tested on" is a sentiment any experimentation leader will recognize. Ashley put the underlying math plainly on the show: most new features lose, which is exactly why measuring matters, but to an individual player who doesn't see that math, testing can sound like a scary thing.
Supercell's answer is to treat the fear as legitimate and design around it. Announce the test. Explain the uncertainty. Guarantee the make-up. The result is an experimentation program that a skeptical, passionate community tolerates, and even participates in, because the fairness contract is explicit.
Everyone can analyze an experiment now
The third act of the conversation is the newest, and it will sound familiar to any team watching AI reshape their workflow.
"It's a really new era," Shan said. "We have some experiment skills they can just import to Claude, and then they can ask Claude, okay, use this skill and now I have been set up this experiment and data. You can find the data yourself and analyze it. Tell me the result. What shall I do?"
The consequence is a wholesale change in who does analysis. "It's not restrict to only data analyst or product manager with experiment knowledge. Basically everyone can run and analyze experiment on their own now."
For a central team whose mission is spreading experimentation culture through a decentralized company, this is the dream scenario. Analysis used to queue behind a small number of qualified people. Now it doesn't. But Shan is candid about what got traded away. "It also creates a new challenge that we don't know the quality of running those AI-driven analysis. I'm sure that people with experience, they can guide the AI to run the analysis correctly, but we don't know everybody who's running analysis using AI."
When quality lived inside the gatekeepers, removing the gate meant removing the guarantee. And AI's failure mode is uniquely dangerous in this domain, as Ashley noted: it can give you a very confident answer built on half the data. A false winner doesn't just waste one launch. It quietly redirects investment and, eventually, undermines trust in the whole program. The emerging answer, the one GrowthBook is investing in, is to move rigor from people into the system itself: experiment templates, guardrails, and defaults that make the AI behave the way your best data scientist would.
Trust is the infrastructure
Pull the three threads together and Supercell's approach resolves into a single principle: experimentation at scale runs on trust, and trust has to be engineered as deliberately as the statistics.
The game teams trust the central team because it partners and shares learnings instead of dictating. The creative culture trusts experimentation because hypotheses serve the craft rather than replacing it. Players trust the tests because the company announces them and guarantees fairness. And the next frontier, AI-driven self-serve analysis, will earn trust the same way, through guardrails that make every analysis as rigorous as the expert-run ones.
Fewer than a hundred tests a quarter for 300 million players sounds like restraint. It is actually the cost of doing every one of them in a way nobody, inside the company or out, has reason to doubt.
Ready to bring that kind of rigor to your own experimentation program? Learn more at growthbook.io.
Related articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics — free.


.avif)

.avif)