All books
Jacket for Proof Before Production — Evaluation systems that can stop a release

Evaluation & Reliability

Proof Before Production

Evaluation systems that can stop a release

A probabilistic system cannot be shipped on taste, so evaluation has to be built like a build system: versioned datasets, trajectory scoring, judges you have tested, and gates wired into CI.

The demo passes because someone chose the inputs. The product fails quietly because production did not. This book builds the apparatus that closes the gap: a definition of done for open-ended output, a first dataset drawn from real traffic, metrics separated into outcome, trajectory and system, a judge calibrated against human annotation that carries its own error bar, and a suite with the authority to fail a deploy. It also draws the line where the OpenTelemetry GenAI conventions stop and your own evaluation data has to start.

What it makes operable

  1. 01Define done for an open-ended output before writing a single eval
  2. 02Build the first dataset out of real traffic and keep the seed set uncontaminated
  3. 03Score trajectories, then calibrate the judge against a human-annotated reference set
  4. 04Wire the suite into CI as a gate with the authority to fail a deploy

Contents

12 of 12 published
  1. 01Why Demos Pass and Products FailThe distribution a demo samples, and the one production actually sends.10 min
  2. 02Defining Done for an Open-Ended OutputTurning "a good answer" into criteria two reviewers would agree on.10 min
  3. 03Building the First Dataset From Real TrafficSampling, labelling and splitting a seed set you will not contaminate.10 min
  4. 04Outcome, Trajectory, and System MetricsThree families of measurement, and what each one cannot tell you.11 min
  5. 05Scoring the Path, Not Just the AnswerPath-invariant assertions for runs that legitimately take different routes.10 min
  6. 06LLM Judges Need Their Own Test SetsThe judge as a system under test, against a human-annotated reference set.10 min
  7. 07Calibration, Bias, and Disagreement With HumansChance-corrected agreement, bias probes, and a published error bar per score.10 min
  8. 08Benchmarks Saturate; Your Suite Should NotWhy public scores stopped discriminating, and what to grow in their place.10 min
  9. 09Wiring Evals Into CI as a Release GateThresholds, regression rules, and who is allowed to override a red run.10 min
  10. 10Online Evaluation and Guarded RolloutsShadow traffic, canaries, and the signal that triggers a rollback.10 min
  11. 11Tracing With OpenTelemetry GenAI ConventionsThe span hierarchy the conventions define, and where they deliberately stop.10 min
  12. 12Incident Response for Systems That Are Fluently WrongDetecting, containing and explaining an error that reads as confident.11 min

Evidence

4 sources

Third-party sources behind the book's premise. Every figure in them belongs to the party that published it and is attributed to them in the text.