
Notable Quotes
"Unit tests, like traditional software unit tests don't work with AI."
"Whenever we have any kind of like question dilemma, we just put it to the test. We just run an A/B test."
"This is so counterintuitive that my manager forced me to run a confirmatory experiment, check the data, make sure that there's no sort of a selection bias."
"AI development is inherently very uncertain very experimental, because you never know how much ground you have covered."
"Decision-making is always gonna remain fundamental, and decision-making can only essentially be done through experiments."
Takeaways

Fin A/B tests everything, even one-character prompt changes, and treats a 20 to 30 percent win rate as a sign of a healthy program.

A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.

More conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations.

