Evals

Grade any model against your own test cases through your real Gateway setup

Evals grade a model against your own test cases and return a pass or fail verdict. Use them before adopting a model, to gate a migration, or on a schedule for regressions. Runs take the production path, so your credentials, routing restrictions, and guardrails apply, and every call appears in your logs and spend.

Suites and test cases

Create suites on the Evals page. Each test case has input messages (a prompt or multi-turn conversation), one or more graders, and a Required flag (on by default). Add cases in the dashboard, import JSONL or CSV, or choose Save as test case on a request in Logs. To keep cases in your repo, sync them over the API.

Graders

Prefer deterministic graders: they’re exact and free. The LLM judge costs an extra model call per case, so save it for tone or accuracy.

GraderPasses when
Exact matchThe response equals an expected string
Contains / does not containA substring is present or absent
RegexThe response matches a pattern
JSON schemaThe response is JSON that validates against a schema
JSON fieldA field at a path such as result.items[0].sku passes a comparison (eq, ne, gt, gte, lt, lte, contains, exists)
Tool callA named tool was called (or not), optionally with argument assertions
LLM judgeA judge model scores the response 0 to 1 against a rubric and meets the threshold

A case passes when all its graders pass (or any, if you choose any); grader errors count as failures. Set the default judge model, rubric, and threshold on the suite’s Configuration tab and override them per case. Judge calls bill to your organization.

Pass criteria

Set pass criteria on the Configuration tab. Errored cases fail under both options.

CriteriaThe suite passes when
All enabled cases must pass (default)Every enabled case passes
A minimum share of cases must passThe pass rate across enabled cases meets your percentage and every required case passes

The verdict is advisory: it informs a migration cutover but never blocks one.

Reference

LimitValue
Cases per suite500
Graders per case20
Input messages per case50, up to 32 KB total
Import file size1 MB
Per-case request paramstemperature, max_output_tokens, response_format, tools, tool_choice

Next steps