Runs and monitoring

Run suites with repeated trials, read verdicts with confidence intervals, and get alerted on failures

A run grades every enabled case in a suite against one model. Choose Run evals on the suite page, review the estimated cases, judge calls, and spend, and confirm. Runs continue if you leave; a suite can have 3 in progress.

Trials

Each case runs 1, 3, or 5 times and passes only if every trial passes. Use 1 while iterating and 3 or 5 for real decisions; cost scales with trials.

Reading a verdict

Pass rates show a 95% confidence interval: 8/10 reads “80%, CI 49-94%”, which is weak evidence. Runs with fewer than 20 graded cases are flagged Low evidence, with a hint of how many cases would tighten the interval.

A case result shows the full response, each grader’s outcome (with the judge’s score and rationale), latency, and cost. Each run is a trace with one span per case.

Comparing two runs

Run the suite against both models, then choose Compare on either completed run. Gateway pairs results by case, counts cases fixed and broken, tests significance, and lists regressions first.

Alerts

Alerts appear at Evals → Alerts, with an open count in the Evals header. Acknowledging records who and when. Cancelled runs never alert.

AlertRaised when
Evals failedA run completed without meeting the pass criteria
Run errorA run couldn’t execute, for example because every case errored
RegressionA scheduled run passed, but more cases regressed than improved against the previous completed run of the same suite and model. Errored cases don’t count.

Route alerts to Slack or a pager with a webhook on the alerts page.

Scheduled runs

On the Configuration tab (or over the API), schedule a suite against one model every 6 hours, 12 hours, day, or week. Saving restarts the clock, so a re-enabled schedule never fires a backlog; a run due while another is in progress is skipped.

Next steps