Why GrowthBook’s MCP server is just the interface

What GrowthBook learned by rebuilding its agent layer around APIs, skills, and human judgment.
The easiest way to misunderstand GrowthBook’s MCP server is to treat it as the headline.
An MCP server built on the Model Context Protocol can make for a neat launch story: take a product API, expose a set of tools, add a configuration snippet, and let an AI agent do something that used to require a dashboard.
Useful, yes. But increasingly ordinary.
The more important question is what the product has to know, expose, validate, and remember for any of those tool calls to be useful.
In a conversation with Alyssa Nicoll and Scott Bailey, the story kept pulling in two directions. Alyssa kept widening the frame beyond MCP toward GrowthBook’s larger AI strategy. Scott kept tracing the system downward, from skills to the CLI to a deep API. Those directions meet at the same conclusion: MCP is not the product. It is one way into the product.
GrowthBook’s server matters because the company decided the protocol layer should get thinner while the system behind it got deeper.
The first server solved an interface problem
When GrowthBook introduced its first MCP server, the immediate problem was obvious.
Software development was moving into agent surfaces such as code editors and terminal-based assistants, while feature flagging still lived partly in a separate interface. If an agent could write the feature but could not create the flag, inspect the experiment, or understand the release context, the workflow stopped halfway.
“How people write software was changing,” Scott said, “and feature flagging in particular should be right where people are already writing their code.”
GrowthBook described that first release as the first open-source, production-ready MCP server for experimentation and feature management. The original architecture fit the moment: it exposed a collection of named tools for concrete jobs such as creating flags, reading experiments, generating types, and searching documentation. Before reusable agent skills became a common design pattern, the MCP tool list carried much of the product knowledge.
The expected user was mostly a product engineer. They were already in an editor, already working on a change, and needed GrowthBook to be available in the same conversation. The value was proximity. The agent did not have to send the engineer back to a browser just to finish the release workflow.
That was the correct bet for the first version, but was not the final architecture.
Then the maintenance bill arrived
There was no dramatic failure. No rogue agent shipped a broken experiment. No launch had to be rolled back. The lesson was more mundane and more useful: GrowthBook had accumulated too many surfaces that could drift.
The REST API, CLI, MCP server, and skills were all reasonable on their own. Together, they created a maintenance problem. A product change could require updates in several places. A new endpoint did not automatically make the MCP layer smarter. A better workflow explanation could end up trapped in one tool description. For a scaling engineering team, every duplicated surface became another place where the public promise could separate from the product.
Rather than keep expanding a parallel catalog of MCP-specific tools, GrowthBook rebuilt v2 as a thin routing layer over the API and skills it already maintains.
The current thin server exposes four tools by name: growthbook\_list\_skills, growthbook\_read\_skill, growthbook\_api\_read, and growthbook\_api\_write. The current open-source skill repository provides around 30 skills, with a small set of high-level skills routing to more than two dozen workflow skills. The MCP server loads those workflow skills on demand instead of expanding into a tool for each endpoint, while the read/write split gives clients a clearer signal about which calls may change state.
The maintained skills become the authority on how to use GrowthBook well. The API remains the authority on what the product can do. The MCP server connects an agent to both.
“Every time we update a skill, the MCP server gets better,” Scott said. “Every time we update the REST API, the CLI stays in sync with it.”
Capability and judgment belong in different layers
The cleanest way to understand the new architecture is to separate capability from judgment.
The API supplies capability. GrowthBook has been expanding a typed REST API across feature flags, experiments, metrics, data sources, SDK connections, and product configuration. Its newer endpoints are defined from shared schemas that support runtime validation and OpenAPI generation. The CLI is generated from that API definition, which helps it stay current as the product changes.
Skills supply judgment. An API endpoint can tell an agent how to create an experiment. A skill can tell the agent when an experiment is warranted, what context to inspect first, how to choose metrics, which assumptions need human confirmation, and what evidence to return for review. The endpoint is a verb. The skill is a playbook.
That split is visible in an actual read-only path. For “What should we test next?”, the client does not search a giant catalog for a bespoke brainstorm tool. It discovers a domain skill, loads the relevant workflow skill, and only then calls the API:
growthbook_list_skills({})
growthbook_read_skill({"name":"experiments"})
growthbook_read_skill({"name":"experiments/references/experiment-brainstorm"})
growthbook_api_read({"path":"/api/v1/experiments?limit=50&status=stopped&sortBy=dateCreated&sortOrder=desc"})The first three calls load judgment into the conversation. The fourth fetches company data through a read-only tool. If the user later authorizes a mutating step, the agent moves to growthbook\_api\_write, whose schema accepts POST, PUT, PATCH, or DELETE and marks the operation as potentially destructive. The MCP layer identifies the kind of action; the skill supplies the procedure; GrowthBook permissions and review state still decide what the action can change.
That distinction changes what the team maintains. Instead of teaching the same workflow separately to every AI surface, GrowthBook can improve a canonical skill and make that guidance available through the MCP server, direct agent use, and the in-app assistant. Instead of manually mirroring every API change in a CLI, it can regenerate the CLI from the same product contract.
It also prevents an important category error: a skill is not a security boundary. Skills are maintained instructions, and an agent can still misunderstand or apply instructions imperfectly. Permissions, validation, review states, approval policies, and audit history belong in the product layer. Tool annotations can help an agent distinguish reads from writes, but they do not replace scoped credentials or product governance.
The result is a stack in which each layer has a clearer job. The API makes the product addressable. The CLI and MCP server make it reachable from different agent environments. Skills make common workflows legible. GrowthBook itself remains the system that holds state, evaluates rules, records experiments, and applies organizational controls.
Different agent surfaces can draw on the same maintained workflows and product capabilities.
- MCP client
- Agent with a CLI
- In-app assistant
What to inspect, which procedure to follow, what assumptions to check, and what evidence to return.
The operations available for flags, experiments, metrics, and configuration, backed by a typed contract.
Permissions, validation, review states, and audit history govern what can change. Experiment history and curated learnings supply context.
Skills guide an agent; they do not grant authority. The controls available depend on the team’s plan and configuration.
Every experiment becomes context for the next one
Any team that runs experiments in GrowthBook is already building the raw material for this workflow. The team’s experiments, configurations, metrics, and results remain available as part of its working history. The customer Scott described did not need to create a special AI corpus. It connected an agent to context the team had accumulated by using GrowthBook.
Since that conversation, GrowthBook has introduced Learnings: a place to capture conclusions across experiments and other evidence, including findings that support or contradict a pattern. Teams can make those curated learnings available to agents through the API and MCP, supplementing the experiment-brainstorm skill’s review of individual tests.
That makes the leap smaller than it sounds. Instead of starting from a blank prompt, an agent can inspect what the team has tried, which questions it has already answered, and where a follow-up might go deeper.
In the anonymized example from the interview, the customer connected its in-house agent through GrowthBook’s MCP server and asked for experiment ideas grounded in that existing history. The experiment-brainstorm skill guided the agent to inspect prior work, propose extensions or unexplored directions, and carry a selected idea into experiment design.
The actual experiment-brainstorm skill makes “grounded in history” testable. It reads up to 50 stopped experiments, starting with the newest, fetches results for roughly 20 of them, checks winners, losers, inconclusive tests, sample ratio mismatch flags, and under-explored projects or tags, then requires every proposal to cite a past experiment by name or ID. It also forbids creating an experiment during brainstorming. The person reviewing the shortlist chooses which idea to develop.
From there, the agent could do the heavy setup work: draft the experiment, connect or create the underlying feature flag, configure targeting, and prepare the object for launch. The same lifecycle continues after exposure begins, with skills organized around analysis and stopping decisions.
The interesting part is not that the agent can call several endpoints in sequence. It is that the sequence reflects the way an experimentation team actually works: brainstorm, design, launch, analyze, decide. The API provides the atomic operations; the skill turns them into a coherent job.
The same operating model shows up at scale. In GrowthBook’s Fyxer case study, a four-person growth team used GrowthBook and AI-assisted workflows to run 541 A/B tests in a year. The case study says the MCP server helped automate the path from Cursor to GrowthBook and into Slack. About 25% of those experiments were winners, while Fyxer grew from \$1M to \$35M ARR. That is not evidence that MCP caused the growth, and the case study does not isolate time saved by the server. It does show the operating model at meaningful scale: agents compress the work around a test; the experiment still decides what earns rollout.
Build on your experiment history
Explore the experiment-brainstorm workflow on GitHub, then ask your connected agent for ideas grounded in your past tests. Review the shortlist before designing a new experiment.
View Brainstorm WorkflowAutonomy is a risk decision, not a product setting
AI conversations often turn autonomy into a personality test. Either you run in “YOLO mode,” or you demand a human click for every action. Neither describes how most organizations actually adopt agentic work.
Alyssa joked that her instinct was to give an agent broad permission and ask it to stop interrupting her. Her more useful answer was about the placement of judgment.
“The goal isn’t human approval on every step or every API call,” she said. “It’s putting approval at the consequential decision points along the way.”
Consequential does not simply mean “write operation.” Brainstorming ten experiment ideas is usually cheap and reversible. Drafting a flag or experiment may also be low risk when nothing is exposed. Launching a targeting rule changes a customer experience. Stopping an experiment or acting on an analysis can redirect product work even if the underlying data was read-only.
The right boundary therefore depends on the decision, the environment, the team’s operating model, and its confidence in the surrounding controls. One company may let an agent configure a complete draft but require a PM or engineer to approve launch. Another may allow scoped writes in a sandbox. A mature team may automate more of a familiar workflow while still reserving unfamiliar or high-blast-radius changes for review.
A team could allow inspection and scoped draft work, then require approval before a change reaches customers.
Read prior experiments and propose follow-ups.
Example boundaryUse scoped read access. A person chooses which idea to develop.
Return: a shortlist with cited experiments.
Draft the experiment, flag, targeting, and metrics.
Example boundaryAllow only authorized draft or sandbox writes; keep exposure unchanged.
Return: proposed objects and assumptions.
Launch targeting or expand a rollout.
Example boundaryRequire a named reviewer, agreed guardrails, and a rollback owner.
Return: the exact change awaiting approval.
A read can still lead to a consequential decision. Stopping a test or acting on an analysis also deserves a review boundary chosen by the team. Read/write labels alone do not settle the policy.
GrowthBook’s recent governance work is relevant because it applies product controls to actions regardless of whether they came from a human in the UI or an agent through an API. Draft and review states, approval policies, permissions, validation, and audit trails can keep agent work visible and inspectable. The exact controls available still depend on the organization’s configuration and plan, so they should be verified rather than assumed.
The design goal is not zero human involvement. It is a smaller, more valuable human role: decide the policy, inspect the evidence, and approve the moments that change exposure or interpretation.
Why letting an AI act still feels risky
The fear is not abstract. Plenty of people are comfortable asking AI to explain code and deeply uncomfortable letting it change anything on their behalf. The memes make the anxiety funny—agents inventing facts, deleting the wrong thing, or sending the message—but the concern underneath is reasonable. Once AI can act, a bad answer can become a bad change.
Some of that hesitation is personal, too. People wonder whether teaching an agent to do more of their job is the first step toward making their own judgment less valuable. Scott’s argument was more practical: an expert who can use an agent to run more experiments and create more value is not removing expertise from the process. They are extending its reach.
Scott did not dismiss the risk or answer it with a blanket claim about model quality. He started with context. An agent working inside a codebase can see how the application already initializes GrowthBook, how flags are evaluated, how tests are written, and where event tracking happens. GrowthBook can add the other half: projects, environments, flags, experiment history, metrics, and workflow guidance. Better context does not make an agent infallible, but it reduces how much it must guess.
Trust then depends on four practical conditions:
- The agent can see the relevant application and GrowthBook context before it acts.
- Its credentials and tools limit the authority it does not need.
- The team can inspect proposed changes before they take effect and knows how to reverse them.
- The team can repeat a low-risk workflow and compare the result with its own judgment.
“Build trust by doing,” Scott said.
That is a more credible promise than saying agents will not make mistakes. The point of governance is not to prove that a model cannot fail. It is to make mistakes easier to catch, constrain their blast radius, and preserve a record of what changed. A team can start with inspection, move to sandbox drafts, and widen authority only after the outputs earn it.
This is also why the agent should return artifacts rather than reassurance: the GrowthBook object, the targeting rule, the code diff, the metric choice, the assumptions it made, and the decision still waiting on a human. Trust grows when the work is legible to someone who was not in the chat.
The audience grew beyond engineers
The first MCP server was naturally developer-centered. The newer architecture supports a broader reality: not everyone uses the same agent surface, and not everyone who needs agent assistance works in a code editor.
Some teams want a CLI because their agent has a terminal. Some use an MCP client where direct shell access is unavailable or undesirable. Some PMs, growth practitioners, and analysts will work through GrowthBook’s in-app assistant. The adapter changes; the product knowledge should not.
“Not everyone uses agent services the same way or uses the same ones,” Scott said. The advantage of a deep API with reusable skills is that GrowthBook can meet those surfaces without inventing a separate product model for each one.
Alyssa put the same point more bluntly: “The way we are supporting developers goes so much beyond the MCP server.”
That broader framing matters. If GrowthBook equated its AI strategy with MCP, it would be betting the product on one protocol and one class of client. By treating MCP as an adapter, it can support the editor, terminal, chat client, and in-app experience while keeping the same experimentation concepts underneath.
It also changes the persona. A PM does not need to become a Cursor power user to ask what an experiment means or prepare a follow-up. An engineer does not need to leave the repository to create the release control surrounding a code change. Both can work through the surface that fits their job, then meet in the same GrowthBook objects and review process.
The real product is what survives the interface
GrowthBook’s MCP story is not really about a protocol win. It is about choosing where the durable parts of an agent experience belong.
If the API contract, experiment history, skills, permission model, review state, and audit trail remain useful when the interface changes, they are product infrastructure. If the intelligence lives only in a tool description, it is a demo with a maintenance plan.
Protocols will change. Agent interfaces will come and go. MCP is the bridge. The product is everything that can cross it.
Related reading
- Build and run charts with the GrowthBook MCP Server
- Keep feature flag code type-safe with the GrowthBook MCP Server
- Launch an A/B test from your IDE with the GrowthBook MCP Server
- Search and audit fact metrics and fact tables with the GrowthBook MCP Server
- Ship a feature behind a flag from your editor with the GrowthBook MCP Server
Put the flag skill to work
Open GrowthBook’s feature-flags skill on GitHub to see the workflows for creating flags, configuring targeting, and preparing rollouts. Use it in your agent with your team’s permissions and review process.
Explore Feature Flag SkillRelated articles
Ready to ship faster?
No credit card required. Start with feature flags, experimentation, and product analytics — free.





.png)
.png)