Evals
Evals grade a model against your own test cases and return a pass or fail verdict. Use them before adopting a model, to gate a migration, or on a schedule for regressions. Runs take the production path, so your credentials, routing restrictions, and guardrails apply, and every call appears in your logs and spend.
Suites and test cases
Create suites on the Evals page. Each test case has input messages (a prompt or multi-turn conversation), one or more graders, and a Required flag (on by default). Add cases in the dashboard, import JSONL or CSV, or choose Save as test case on a request in Logs. To keep cases in your repo, sync them over the API.
Graders
Prefer deterministic graders: they’re exact and free. The LLM judge costs an extra model call per case, so save it for tone or accuracy.
A case passes when all its graders pass (or any, if you choose any); grader errors count as failures. Set the default judge model, rubric, and threshold on the suite’s Configuration tab and override them per case. Judge calls bill to your organization.
Pass criteria
Set pass criteria on the Configuration tab. Errored cases fail under both options.
The verdict is advisory: it informs a migration cutover but never blocks one.