Waymo has logged 220 million fully autonomous miles with 17 times fewer serious crash injuries than human drivers. At VB Transform 2026, director of engineering Manasi Joshi explained the methodology behind that number: a discipline she calls eval-centric development, where evaluation is an engineering requirement from day one, not a pre-launch checklist. The maturity of a project's tests, she said, is how Waymo measures the maturity of the project itself.

The evaluation stack runs continuously across model training, post-training, and both open-loop and closed-loop simulation. Waymo tests against billions of synthetic miles and curates specialized datasets for high-stakes edge cases: vulnerable road users, railroad crossings, construction zones. Critically, no release decision is fully automated. Internal safety leads approve every software push and service-area expansion. Joshi was direct: human lives are at stake, and human oversight is non-negotiable. The company also flags a problem that undermines most AI quality claims: a performance metric is only as trustworthy as the dataset behind it.

The piece is worth reading in full because Joshi also covers Waymo's compute constraints, its shift from transformers in 2017 through to today's generative multimodal and vision-language-action models, and how the company evaluates the AI agents its own engineers use internally. The throughline is a concrete, transferable framework: define a real business outcome, build representative eval data, keep testing after launch, and assign named humans who stay accountable. That framework applies whether the system is navigating a construction zone or handling a customer refund.

[READ ORIGINAL →]