How 12 of the world's top product teams use AI to run experiments

AI is not replacing experimentation. It is increasing the number of ideas teams can build, which makes disciplined experimentation more important.
The useful question is no longer whether product teams use AI. It is where they put AI inside the learning loop.
Some teams use agents to turn a hypothesis into code. Others use AI to search prior research, write SQL, or summarize an experiment. Teams building AI products use controlled experiments to compare prompts, models, retrieval strategies, and user experiences that offline evaluations cannot settle. The most mature teams combine all three patterns while keeping people responsible for the decisions that carry product risk.
That distinction matters. A model can produce a plausible hypothesis, a clean implementation, and a confident readout while still optimizing the wrong outcome. The examples below show a more grounded approach: let AI compress mechanical work, then use controlled experiments to determine whether the resulting change actually helps users.
Three ways AI enters the experimentation loop
Across these 12 teams, AI plays 3 different roles. Treating them as one use case obscures the controls each one needs.
| Role for AI | What it does | Main risk | Human checkpoint |
|---|---|---|---|
| Researcher or analyst | Synthesizes calls, finds prior tests, queries data, and drafts readouts | Confident summaries can hide missing context or weak data | Verify sources, freshness, segments, and data quality |
| Builder or operator | Creates variants, code, flags, and experiment configuration | Fast implementation can amplify a flawed hypothesis or unsafe rollout | Review the diff, metrics, targeting, and launch plan |
| Experimental subject | Supplies the model, prompt, agent, or retrieval system being compared | Offline scores may not predict real user outcomes | Pair evals with controlled production measurement |
The strongest programs do not ask an agent to invent the entire process on every run. They encode reusable templates, metric definitions, and approval steps. The agent handles repeatable operations inside those boundaries.
This resembles the broader lesson from Microsoft's long-running experimentation platform: scale depends on trustworthy shared infrastructure, not merely a larger queue of tests. AI changes the speed and accessibility of the work, but it does not repeal assignment integrity, statistical power, or the need for a control.
AI as researcher and analyst
Research synthesis is often the safest place to begin. AI can process sales calls, support conversations, experiment descriptions, and product data faster than a person can read every artifact. It can also retrieve the history that teams routinely lose across documents and dashboards.
The limitation is provenance. A summary is only useful when the reader can trace it back to the underlying calls, queries, or experiment results. Teams should ask the system to show its evidence, state what it could not access, and distinguish an observed result from an inference.
AI as builder and operator
Coding agents can shorten the distance from an approved idea to a testable variation. They can implement a small interface change, wire a feature flag, prepare an experiment draft, and return a verification receipt. This is where teams can remove days of scheduling and handoffs.
It is also where a weak idea can reach production unusually quickly. The control system has to live outside the model: scoped permissions, reviewed changes, approved metrics, gradual exposure, and an accessible kill switch. The NIST AI Risk Management Framework offers a useful general principle here: organizations still need explicit governance, measurement, and management around AI-enabled decisions.
AI as the thing being tested
For AI products, the model is part of the experimental variable. Teams may compare a prompt, model provider, fine-tuning technique, retrieval configuration, tool set, or complete agent workflow.
Offline evaluations remain essential for catching regressions before exposure. They cannot fully predict whether users will understand, trust, or benefit from a new behavior. Intercom's account of designing AI products recommends production A/B tests when interaction volume supports them, combined with telemetry and qualitative feedback when it does not. That layered testing approach appears repeatedly in the teams below.
Test AI beyond offline evals
Learn how to combine offline evaluation, guarded rollout, and controlled experiments when shipping generative AI.
Compare Evals and A/B Tests12 product teams show where AI is useful
The examples are not 12 copies of the same playbook. Some teams are automating the experimentation lifecycle; others are using experiments to make AI products safer and more useful. Together, they show where AI is producing real operating leverage and where human judgment remains essential.
1. Fyxer turns one experiment brief into a coordinated workflow
Fyxer's small growth team demonstrates the most complete operator pattern. The company reported running 541 A/B tests in a year while growing from $1 million to $35 million in annual recurring revenue. Its workflow connects a structured experiment brief to Claude, Cursor, internal systems, and GrowthBook.
AI does not simply suggest button-copy ideas. Shared Claude skills turn repeatable tasks into team workflows. Cursor helps implement and preview contained changes. An AI data analyst connected to BigQuery answers segmentation questions, while automations prepare documentation and identify stale experiment code. The team can move several bounded experiments forward in parallel because the tools reduce implementation and analysis queues.
The important part of the Fyxer experimentation story is not the test count by itself. It is the structure around the speed: a brief, reusable skills, reviewable changes, shared data context, and an experiment system of record. AI compresses the loop because the loop already has named stages.
2. DoorDash uses AI to make an enormous experiment history usable
DoorDash reportedly processed 12,000 experiments in a year across a 3-sided marketplace. At that scale, institutional memory becomes an infrastructure problem. A person cannot remember what every consumer, merchant, and Dasher experiment taught the company.
The team's AI opportunities include mining prior experiments for useful context, helping people configure and debug tests, identifying sample ratio mismatch problems, and drafting company-wide readouts. Those uses have something in common: AI makes a complicated system more accessible without deciding the business tradeoff on its own.
The DoorDash interview makes the boundary explicit. An AI system can surface research and explain results, but people must decide how to balance consumer price and speed, merchant economics, and Dasher earnings. There is no single metric the model can maximize without judgment.
3. Diligent collapses weeks of research synthesis into hours
Dan Layfield's experience across Diligent, Codecademy, and Uber Eats points to a near-term use of AI that requires little speculation: synthesis. Instead of manually scheduling, transcribing, and coding a large set of customer conversations, his workflow uses AI-accessible call data to identify patterns quickly. He also uses Claude as a practical assistant for straightforward experiment analysis.
This does not remove research judgment. It changes the cost of reaching the point where judgment is possible. A product manager can inspect themes while the evidence is still timely, compare those themes with product behavior, and decide whether a losing test exposed a bad idea or an incomplete implementation.
The Diligent conversation also shows the danger of treating a first inconclusive result as a final answer. AI can accelerate analysis, but it should not automate the decision to abandon a strategically important hypothesis.
4. Kargo uses AI to widen access to technical evidence
At Kargo, AI's immediate value is access. People who could not previously write SQL or navigate a complex ad-tech system can ask better questions of the data and participate more directly in experimentation. That broadens the set of people able to form and challenge hypotheses.
Kargo also illustrates an architectural limit. Its bidding systems operate under severe latency constraints, so large language models do not belong in the live auction path. The team can use agents around the system for orchestration, investigation, and test operations without pretending an LLM is appropriate for every runtime decision.
That nuance is the best part of the Kargo experimentation story. “Use AI” is not an architecture. Teams need to identify where flexible reasoning helps and where deterministic, low-latency software must remain in control.
5. Box applies AI to the next experiment, not the final decision
Box's ecommerce team already treats experimentation as a revenue discipline. Its pricing-page tests showed that a general direction such as “simplify” can win once and then lose when taken too far. That history is valuable input for AI-assisted ideation because it contains the boundary conditions a generic model lacks.
The team is exploring agents that scan competitive changes and help generate experiment ideas. The safe formulation is help generate. A competitive change is not proof that Box should copy it, and a generated hypothesis is not a reason to ship. The experiment still has to isolate the change and measure whether it helps Box's customers.
The Box interview suggests a productive sequence: use AI to expand the idea set, use prior company evidence to narrow it, and use controlled experiments to adjudicate the remaining uncertainty.
6. JPMorgan Chase treats AI velocity as a demand for more measurement
JPMorgan Chase offers a useful counterweight to stories about autonomous experimentation. Kevin Yang's team supports roughly 100 product teams and has grown the program from 8 experiments in its first year to around 300 annually. His view of AI is not that it makes controlled measurement obsolete. It makes measurement more urgent.
When more people can build and customize features, more changes reach the point where they need causal evidence. Without controls, errors also compound faster. This is especially important in financial products, where removing friction may increase one engagement metric while weakening trust or customer protection.
The JPMorgan Chase discussion therefore puts AI inside an established decision system: self-service infrastructure, experts supporting consequential initiatives, agreed metrics, and a plan for what the team will do if a test loses.
7. Fin uses production experiments as the gold standard for AI behavior
Fin, Intercom's AI support agent, represents AI as the experimental subject. Offline evaluation can cover thousands of prepared examples, but live users create a far larger and less predictable distribution of questions. Fin therefore uses controlled production experiments to determine whether prompt, context, latency, and product changes improve actual outcomes.
One experiment found that increasing latency could improve positive feedback, possibly because an instantaneous answer felt canned. Another exposed a serious failure: giving the agent more conversation context improved some metrics while increasing unauthorized promises about refunds. The team revised the behavior and tested again rather than shipping the aggregate winner unchanged.
Intercom has separately described running more than 120 A/B tests while developing a major Fin release. Its Fin 2 launch account reinforces the lesson from the GrowthBook interview with Fin: AI quality is multidimensional, and a headline metric cannot substitute for guardrails and failure analysis.
8. Khan Academy measures whether AI helps students learn
Khan Academy cannot evaluate an AI tutor by asking only whether students send more messages. Engagement may mean productive struggle, but it may also mean confusion. The team needs metrics closer to learning quality, correctness, and responsible tutoring behavior.
Its approach combines AI-driven evaluation with controlled production experiments on Khanmigo. The Khan Academy experimentation session describes governance across research, analytics, engineering, and product, rather than treating model evaluation as a standalone ML task.
This is a pattern other AI teams can copy: use evals to screen candidate behavior, then measure the outcomes only real users can produce. The system should know whether an answer is acceptable before rollout and whether the experience actually helps after rollout.
9. Character.AI compares model techniques from the user's perspective
Character.AI uses controlled experimentation to connect post-training decisions with product outcomes. A modeling technique may improve an offline benchmark and still make conversations less satisfying, less engaging, or less trustworthy for a particular user segment.
The operative phrase in Character.AI's account is comparison “from the perspective of our users.” That moves the decision away from model scores alone and toward observed behavior under controlled assignment. It also encourages config-versus-config iteration after the initial AI system is already in production.
The Character.AI example shows why AI experimentation needs both model telemetry and product metrics. Token cost, latency, and an evaluator score explain the system. Retention, successful conversations, and segment effects explain whether the system serves users.
10. Dropbox uses staged exposure for fast-moving AI products
Dropbox's AI products increase the need for rapid iteration, but its scale and security requirements make uncontrolled rollout unacceptable. The company standardized staged exposure through feature rings that move changes from internal users to employees, alpha users, and broader customer populations.
That release structure makes AI experimentation safer because teams can observe behavior and reverse course before full exposure. Dropbox also analyzes experiments against its existing Databricks data lake, avoiding a separate measurement reality for AI launches.
The Dropbox customer story is a reminder that “using AI to experiment” does not always mean allowing an agent to launch a test. Sometimes it means building the release and measurement infrastructure that lets AI teams iterate quickly without skipping control groups, data governance, or rollback stages.
11. Atlassian turns prior experiment knowledge into an agent input
Atlassian's AI Product Builders Week produced a Growth Experiments Agent designed to answer a deceptively valuable question: has someone already tested this, and what did the company learn? That turns experiment history from an archive into an active input to product planning.
The company also emphasizes hands-on use, shared demonstrations, and learning from failed prototypes. Its AI Product Builders Week generated more than 100 documented AI use cases and a library of internal explainers. This cultural layer matters because an agent is only as useful as the context teams make available to it.
The resulting pattern is not “AI creates better hypotheses automatically.” It is “AI lowers the cost of retrieving relevant evidence before a team commits to another build.” That is a smaller claim and a much more useful one.
12. GrowthBook encodes experiment procedures as reusable agent skills
GrowthBook's own product team approaches AI as an operator constrained by an experimentation platform. Open-source agent skills describe repeatable workflows for discovering prior tests, designing an experiment, creating a draft, checking results, and cleaning up temporary flag logic. The MCP Server or REST API supplies capability; the skills supply procedure.
This separation is important. A coding model should not have to rediscover the organization's rules from a prompt. The agent-assisted experimentation workflow keeps templates, metrics, permissions, approvals, and audit history in GrowthBook while allowing the agent to operate from an editor or other supported client.
The practical result is not a fully autonomous product manager. It is a shorter path between an approved decision and verified execution, with a receipt for what the agent read and changed.
What these teams do not delegate
The examples differ, but their boundaries are remarkably consistent.
First, people still choose the problem. AI can cluster feedback and retrieve related tests, but it cannot decide which customer pain or strategic opportunity deserves the team's limited attention.
Second, people define what success means. A model will happily optimize the easiest measurable proxy. JPMorgan Chase's payment friction, Khan Academy's learning quality, DoorDash's 3-sided marketplace, and Fin's hallucination risk all show why the metric set is a business decision rather than a generated default.
Third, people approve exposure. A test that changes a price, financial flow, model behavior, or customer promise deserves an explicit launch boundary. Small initial cohorts, guardrail metrics, and rollback paths make fast iteration compatible with responsible operation.
Finally, people interpret tradeoffs. A primary metric can rise while latency, trust, cost, or a vulnerable segment gets worse. The agent can surface the conflict. It should not decide whose outcome matters.
Google Cloud's discussion of experimentation as an organizational capability reaches a similar conclusion: technology helps teams test and learn, but leaders create the conditions in which learning is valued and acted upon.
Start where your loop is actually slow
The wrong starting point is “automate experimentation.” That instruction is too broad to verify and too vague to govern.
Map the current path from question to decision and record the waiting time at each stage:
- Research: How long does it take to synthesize customer evidence and find related tests?
- Design: How long does metric selection, power planning, and stakeholder review take?
- Implementation: How long does a contained variant wait for engineering?
- Launch: How long does QA, approval, and rollout configuration take?
- Analysis: How long does a trustworthy readout wait for data or analyst capacity?
- Learning: Can another team retrieve the result six months later?
Choose one bottleneck with a clear input and output. Give AI read access before write access. Require linked evidence in generated research or analysis. For implementation work, require a code diff, tests, preview, and flag state. For experiment operations, create drafts first and verify every mutation against the system of record.
Then measure the automation itself. Useful operational metrics include time from approved brief to launch-ready draft, percentage of generated analyses requiring material correction, number of experiments reusing prior evidence, and review time saved. A faster process that creates more cleanup or weaker decisions is not an improvement.
The teams above do not share one AI stack. They share a discipline: AI removes friction from a defined learning system, and experiments keep AI-enabled speed connected to customer outcomes. That is the durable pattern. Build the loop first, place AI where the loop is slow, and keep consequential judgment with the people accountable for the product.
Build more trustworthy tests
Watch Ronny Kohavi and Luke Sonnet explain the design, metric, and review practices that keep high-velocity experimentation credible.
Watch the Experiment SessionRelated Articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics—free.


