An eval suite that doesn't block a release isn't a gate, it's a report. No evals, no ship.

Everyone agrees evaluation matters and almost nobody wires it to a decision. The suite runs, produces a dashboard, someone glances at it, and the change ships regardless. At that point evaluation is documentation of what happened — informative, and structurally identical to having none, because nothing was prevented.
A gate is different in exactly one respect: it can say no. The suite is versioned, runs in CI, and a regression fails the build the way a broken test does. That single property converts a metric into a control, and it is the property that gets negotiated away first, because a blocking check is inconvenient precisely when you most want to ship. The honest question for any team is not "do we evaluate?" but "when did evaluation last stop something?"
Building the gate needs two datasets doing different jobs. A golden set of cases with known-good answers measures whether you are getting better. A regression set of previously-fixed failures measures whether you are breaking what already worked — and it is the one usually missing, which is why systems improve on the demo and quietly degrade at the edges. Without it you are measuring progress while blind to damage.
The best-designed suites also score the path, not just the answer. Trajectory accuracy — did the agent take a sensible route to the result — catches the case where the output happened to be right for the wrong reasons, which is the failure that generalises worst. Judging only final answers rewards luck; judging the route rewards method. Both are checkable, and neither is checkable against a vague standard, which puts the whole apparatus back on the same floor: rules with a pass/fail answer.