Runs and monitoring
A run grades every enabled case in a suite against one model. Choose Run evals on the suite page, review the estimated cases, judge calls, and spend, and confirm. Runs continue if you leave; a suite can have 3 in progress.
Trials
Each case runs 1, 3, or 5 times and passes only if every trial passes. Use 1 while iterating and 3 or 5 for real decisions; cost scales with trials.
Reading a verdict
Pass rates show a 95% confidence interval: 8/10 reads “80%, CI 49-94%”, which is weak evidence. Runs with fewer than 20 graded cases are flagged Low evidence, with a hint of how many cases would tighten the interval.
A case result shows the full response, each grader’s outcome (with the judge’s score and rationale), latency, and cost. Each run is a trace with one span per case.
Comparing two runs
Run the suite against both models, then choose Compare on either completed run. Gateway pairs results by case, counts cases fixed and broken, tests significance, and lists regressions first.
Alerts
Alerts appear at Evals → Alerts, with an open count in the Evals header. Acknowledging records who and when. Cancelled runs never alert.
Route alerts to Slack or a pager with a webhook on the alerts page.
Scheduled runs
On the Configuration tab (or over the API), schedule a suite against one model every 6 hours, 12 hours, day, or week. Saving restarts the clock, so a re-enabled schedule never fires a backlog; a run due while another is in progress is skipped.