Jacket for Proof Before Production — Evaluation systems that can stop a release
Evaluation & ReliabilityFree to read · Open access · 12 of 12 chapters

Proof Before Production

Evaluation systems that can stop a release

A probabilistic system cannot be shipped on taste, so evaluation has to be built like a build system: versioned datasets, trajectory scoring, judges you have tested, and gates wired into CI.

The demo passes because someone chose the inputs. The product fails quietly because production did not. This book builds the apparatus that closes the gap: a definition of done for open-ended output, a first dataset drawn from real traffic, metrics separated into outcome, trajectory and system, a judge calibrated against human annotation that carries its own error bar, and a suite with the authority to fail a deploy. It also draws the line where the OpenTelemetry GenAI conventions stop and your own evaluation data has to start.

What it makes operable

  1. 01

    Define done for an open-ended output before writing a single eval

  2. 02

    Build the first dataset out of real traffic and keep the seed set uncontaminated

  3. 03

    Score trajectories, then calibrate the judge against a human-annotated reference set

  4. 04

    Wire the suite into CI as a gate with the authority to fail a deploy

Contents

12 of 12 published

Every chapter is free to read in the browser, cites its own sources, and stands on its own if you came for one decision rather than the whole argument.

  1. 01Why Demos Pass and Products FailThe distribution a demo samples, and the one production actually sends.10 min
  2. 02Defining Done for an Open-Ended OutputTurning "a good answer" into criteria two reviewers would agree on.10 min
  3. 03Building the First Dataset From Real TrafficSampling, labelling and splitting a seed set you will not contaminate.10 min
  4. 04Outcome, Trajectory, and System MetricsThree families of measurement, and what each one cannot tell you.11 min
  5. 05Scoring the Path, Not Just the AnswerPath-invariant assertions for runs that legitimately take different routes.10 min
  6. 06LLM Judges Need Their Own Test SetsThe judge as a system under test, against a human-annotated reference set.10 min
  7. 07Calibration, Bias, and Disagreement With HumansChance-corrected agreement, bias probes, and a published error bar per score.10 min
  8. 08Benchmarks Saturate; Your Suite Should NotWhy public scores stopped discriminating, and what to grow in their place.10 min
  9. 09Wiring Evals Into CI as a Release GateThresholds, regression rules, and who is allowed to override a red run.10 min
  10. 10Online Evaluation and Guarded RolloutsShadow traffic, canaries, and the signal that triggers a rollback.10 min
  11. 11Tracing With OpenTelemetry GenAI ConventionsThe span hierarchy the conventions define, and where they deliberately stop.10 min
  12. 12Incident Response for Systems That Are Fluently WrongDetecting, containing and explaining an error that reads as confident.11 min

In the age of AI

The advantage was never the model. It's knowing what to build with it — and having a team that can actually ship it.

That's the part I help with: finding where AI genuinely makes your business faster, deciding what's worth building, and standing behind it once it's live.

Four offices, one very full passport

Every dot on this map is a conversation I still remember.

World map showing ViitorCloud offices in Ahmedabad, Zürich, Washington D.C. and Port Louis, and the countries where Vishal Rajpurohit has spoken and travelled
Germany
Indonesia
Saudi Arabia
Spain
Japan
Denmark
Turkey
Singapore
Ireland
Czechia
France
Thailand
Sweden
Mexico
Qatar
Italy
South Korea
Poland
United Kingdom
Malaysia
Belgium
Canada
UAE
Netherlands
Vietnam
Norway
Oman
Portugal
Australia
Austria
New York, USA
Chicago, USA
Las Vegas, USA
San Francisco, USA
Los Angeles, USA
Ahmedabad, India — headquarters
Zürich, Switzerland
Washington, D.C., United States
Port Louis, Mauritius
  • Where I've spoken
  • Office
Sixty seconds from the roadQuick lessons and keynote moments — tap to watch