Filter results

Truist on shipping faster and more safely with AI and human-in-the-loop banking

From chatbots to open-world agents at Microsoft: evals, go-live metrics, and copilot velocity

Upwork on AI-Driven Ops at scale
Top takeaways from our favorite conversations

When you struggle to land a result, lead with the story of what the customer did, then bring the numbers.

Persistence pays: four months and three to four rounds of trial-model testing at Codecademy produced a 35% conversion increase.

Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?
.svg.avif)
A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.

Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.

Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.

Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.

Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.

Documenting experiments in a centralized Wiki creates a growth flywheel: Fanatics' Wiki feeds their roadmap with iterations on already-built features, reducing tech dependency and accelerating velocity.

Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.

Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.
.avif)
Friction can increase revenue. Blocking the "view all" grid and forcing a style choice sent shoppers deeper and lifted conversion and revenue, because the extra click added value.

Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.

Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.

A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.

Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.

The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.





