Sampling releases instead of worshipping averages
Averages hide the awkward releases. Sampling a handful of real ships — good and bad — gives assessment conversations something solid to stand on.
Monthly deployment averages flatten weekends, freezes, and emergency patches into a single comforting number. When we assess reliability, we ask for a sample: three routine ships, two rollbacks, and one release everyone still argues about.
For each sample we note batch size, reviewers present, test gaps known at ship time, and who stayed on after the change. Patterns appear faster than they do in a quarterly slide. The sample also keeps conversations respectful — people discuss a shared artefact instead of defending their reputation against an abstract percentage.