Experiments

How 4 of the top financial services companies experiment under regulation

A graphic of a bar chart with an arrow pointing upward.

Regulation does not eliminate product uncertainty. It changes which questions are eligible, who reviews them, and what evidence must survive the decision.

Financial services teams still need to improve applications, servicing, support, payments, fraud controls, and digital journeys. A blanket ban on controlled experimentation can push changes into broad rollout with weaker causal evidence. An unconstrained growth program can expose customers to unfair, misleading, or harmful treatment.

JPMorgan Chase, Upstart, U.S. Bank, and Ford Credit illustrate a middle path: governed self-service, data control, explicit risk tiers, end-to-end measurement, and an audit trail that connects every result to a decision.

Start with an experiment eligibility policy

Classify proposed treatments before design:

TierExampleDefault path
Routine productNavigation, layout, help contentStandard experiment checklist
Financially materialPricing presentation, payment flow, offer timingProduct, finance, legal, and risk review
Model-influencedRanking, fraud, underwriting supportModel governance, fairness, and performance review
Mandatory obligationRequired disclosure, security fix, accessibility remediationImplement; test only compliant variants or rollout safety
Prohibited or unacceptableDiscriminatory, deceptive, or uncontrolled high-harm treatmentDo not launch

The policy should define approval roles, protected or sensitive attributes, minimum sample and retention rules, geographic constraints, exposure limits, stop conditions, and documentation. GrowthBook's custom fields and shareable workflows can attach ownership and compliance metadata to the record.

JPMorgan Chase: Plan for failure before launch

Kevin Yang's experimentation organization supports roughly 100 product teams and has grown from a handful of annual tests to hundreds. The program attributes substantial value to winners, while Yang argues that prevented harmful rollouts may be even more valuable.

In a regulated environment, that requires deciding the negative path before results create pressure. Which guardrail stops the test? Who can disable treatment? What evidence permits a ramp? What happens if the primary metric improves while complaints or another risk indicator worsens?

The JPMorgan Chase conversation also notes that AI increases product-development velocity, making measurement more important. Faster implementation should not compress independent review or human accountability.

Federal Reserve and OCC guidance on model risk management emphasizes validation, governance, and effective challenge. When an experiment changes a model or model-mediated experience, preserve version, validation evidence, input population, and downstream monitoring.

Practice to copy: Predeclare stop, rollback, escalation, and ambiguous-result rules with the same care as success criteria.

Build trust into every result

Review the power, SRM, stopping, and causal checks that help regulated teams avoid acting on a false signal.

Read the Trustworthy Tests Recap

Upstart: Keep sensitive data and methodology controllable

Upstart operates an AI lending marketplace connecting consumers and financial institutions. Its experimentation stack had become fragmented across feature-flagging and in-house tools, slowing analysis and increasing engineering and analytics dependency.

GrowthBook allowed the company to consolidate workflows, preload standardized metrics, enable engineer self-service, and host data in-house. The Upstart customer story reports experiments moving from days to hours while emphasizing data privacy and control.

The architecture matters because regulated decisions must be reproducible. Teams need to inspect assignment, metric logic, model or feature version, and the population represented. The Consumer Financial Protection Bureau has stated that creditors using complex algorithms still must provide specific reasons for adverse actions. A product experiment cannot make an underlying obligation disappear.

Practice to copy: Keep experiment metadata, exposure, warehouse SQL, model version, and decision rationale available for independent review. Minimize sensitive data in the experimentation control plane.

U.S. Bank: Scale self-service through governed defaults

U.S. Bank's experimentation journey shows the organizational challenge of expanding participation without turning a central team into an approval queue or allowing every team to invent its own standards. The solution is governed self-service: reusable templates, common metrics, training, permissions, and escalation for higher-risk tests.

The U.S. Bank self-service account demonstrates that democratization is not the absence of oversight. The central team owns the safe path and makes routine work repeatable.

NIST's Privacy Framework offers a useful way to identify and manage privacy risk across data processing. Translate that governance into experiment-specific controls: approved identifiers, retention, access, purpose limitation, and deletion responsibilities.

Practice to copy: Let product teams operate low-risk templates independently, while policy automatically routes sensitive populations, financial treatments, and model changes to the required reviewers.

Ford Credit: Connect digital treatment to offline outcomes

Ford Credit's digital prequalification journey ends in a vehicle purchase and financing decision that may occur later at a dealership. Geoffrey Bell's team had to connect online exposure to offline outcomes before experiments could speak the language of receivables.

One team expected an early vehicle selector to improve prequalification. It reduced completion. Moving the selector after prequalification later produced gains, showing that timing—not the concept alone—mattered. The Ford Credit story also describes an experiment linked to a two-point increase in downstream close rate.

This measurement creates responsibilities. Identity resolution should use approved governance, the analysis window should reflect the real decision cycle, and teams should report attrition between online assignment and offline observation. A treatment that changes who reaches a dealership may change the composition of observed purchasers.

The Federal Trade Commission's guidance on using AI and algorithms underscores transparency, fairness, accuracy, and accountability—principles that should appear in product and model experiment reviews.

Practice to copy: Define the causal chain from digital experience to the business outcome and preserve every join, exclusion, and time window used to estimate it.

Design the regulated scorecard

Use four layers:

  1. Customer outcome: completion, time to resolution, successful payment, servicing success, or product comprehension.
  2. Business outcome: qualified application, funded account, receivables, retention, loss, cost, or margin.
  3. Risk guardrails: fraud, complaints, defaults, operational failure, model performance, and adverse events.
  4. Fairness and obligation checks: legally reviewed comparisons, disclosure delivery, accessibility, geographic rules, and protected-group analysis where permitted and appropriate.

NIST's AI Risk Management Framework organizes AI risk around govern, map, measure, and manage. Experimentation fits inside that lifecycle as one source of evidence, not as the entire governance program.

Use the randomized unit for primary analysis and avoid conditioning on a treatment-influenced event. For financial metrics with long delays, include validated leading indicators but schedule downstream reanalysis. GrowthBook's metric layer and fact tables can encode approved definitions on warehouse data.

Preserve an audit-ready decision record

Every regulated experiment should make these questions answerable:

  • What decision and customer problem justified the test?
  • Which policy made the treatment eligible?
  • Who approved product, data, legal, compliance, risk, and model aspects?
  • Who was eligible, assigned, exposed, excluded, and analyzed?
  • Which exact code, flag, model, prompt, disclosure, or workflow varied?
  • Which metrics and SQL governed success and harm?
  • Which automated and human quality checks passed?
  • What did the result show, with uncertainty and limitations?
  • Who made the decision, and what rollout or rollback followed?
  • When were temporary flags, access, and data artifacts reviewed or removed?

GrowthBook supports role-based workflows, self-hosted deployments, inspectable queries, and warehouse-native experimentation. Each institution remains responsible for determining the laws and controls applicable to its product and jurisdiction.

The four companies demonstrate that regulation and experimentation are not opposites. The strongest programs use controls to make uncertainty explicit, evidence reviewable, and harmful changes easier to stop before they reach every customer.

Design for long-term decisions

Hear how experienced leaders approach metrics, failure planning, trustworthy setup, and long-term experiment impact.

Watch the Leadership Session

Table of Contents

Related Articles

See All Articles
Experiments

eCommerce experimentation: Insights and takeaways from the top companies

Aug 17, 2026
x
min read
Experiments

We talked to 4 leaders about getting a stuck experimentation team unstuck — here are their top takeaways

Aug 15, 2026
x
min read
Experiments

We talked to 15 experimentation leaders about losing tests — here are their top takeaways

Aug 14, 2026
x
min read

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.