> This page is for Gateway.

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.merge.dev/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.merge.dev/_mcp/server.

# Evals

> Build eval suites from your own test cases, grade them with deterministic assertions and an LLM judge, and use the verdict as evidence for model decisions like migrations.

Evals grade a model against your own test cases and return a pass or fail verdict. Use them before adopting a model, to gate a [migration](/merge-gateway/evals/model-migrations), or on a schedule for regressions. Runs take the production path, so your credentials, routing restrictions, and guardrails apply, and every call appears in your logs and spend.

## Suites and test cases

Create suites on the [Evals](https://gateway.merge.dev/evals) page. Each test case has input messages (a prompt or multi-turn conversation), one or more graders, and a **Required** flag (on by default). Add cases in the dashboard, import JSONL or CSV, or choose **Save as test case** on a request in **Logs**. To keep cases in your repo, [sync them over the API](/merge-gateway/evals/evals-api#dataset-as-code).

## Graders

Prefer deterministic graders: they're exact and free. The LLM judge costs an extra model call per case, so save it for tone or accuracy.

| Grader                      | Passes when                                                                                                                      |
| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| Exact match                 | The response equals an expected string                                                                                           |
| Contains / does not contain | A substring is present or absent                                                                                                 |
| Regex                       | The response matches a pattern                                                                                                   |
| JSON schema                 | The response is JSON that validates against a schema                                                                             |
| JSON field                  | A field at a path such as `result.items[0].sku` passes a comparison (`eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `contains`, `exists`) |
| Tool call                   | A named tool was called (or not), optionally with argument assertions                                                            |
| LLM judge                   | A judge model scores the response 0 to 1 against a rubric and meets the threshold                                                |

A case passes when all its graders pass (or any, if you choose **any**); grader errors count as failures. Set the default judge model, rubric, and threshold on the suite's **Configuration** tab and override them per case. Judge calls bill to your organization.

## Pass criteria

Set pass criteria on the **Configuration** tab. Errored cases fail under both options.

| Criteria                                  | The suite passes when                                                                   |
| ----------------------------------------- | --------------------------------------------------------------------------------------- |
| **All enabled cases must pass** (default) | Every enabled case passes                                                               |
| **A minimum share of cases must pass**    | The pass rate across enabled cases meets your percentage and every required case passes |

The verdict is advisory: it informs a migration cutover but never blocks one.

## Reference

| Limit                   | Value                                                                         |
| ----------------------- | ----------------------------------------------------------------------------- |
| Cases per suite         | 500                                                                           |
| Graders per case        | 20                                                                            |
| Input messages per case | 50, up to 32 KB total                                                         |
| Import file size        | 1 MB                                                                          |
| Per-case request params | `temperature`, `max_output_tokens`, `response_format`, `tools`, `tool_choice` |

## Next steps

#### [Runs and monitoring](/merge-gateway/evals/running-evals)

Trials, confidence intervals, alerts, and schedules

#### [Model migrations](/merge-gateway/evals/model-migrations)

Use a suite as the evidence gate for a model change

#### [Evals API](/merge-gateway/evals/evals-api)

Trigger runs from CI and read verdicts programmatically