Evals API

Run eval suites from CI: sync cases from your repo, trigger runs, poll verdicts

The Evals API runs eval suites from CI: sync cases from your repo, run them against a catalog model or grade your pipeline’s own outputs, and fail the build on a red suite. See the Evals guide for concepts and graders, and running evals from CI for a walkthrough.

Base URL and authentication

https://api-gateway.merge.dev

Authenticate with the mg_ API key you send /v1/responses with. The organization comes from the key; customer API keys are rejected.

Authorization: Bearer mg_...

Reference

EndpointsWhat they do
GET, PUT /v1/evals/suites, GET /v1/evals/suites/{suite_id}List suites, or create or update one by name (201 created, 200 updated)
PUT /v1/evals/suites/{suite_id}/casesSync cases by external_id. With prune (default true), synced cases missing from the payload are deleted. Dashboard-authored cases are never touched. Returns created, updated, deleted, and unchanged counts.
GET, POST /v1/evals/suites/{suite_id}/runsList or trigger runs
GET /v1/evals/runs/{run_id}, GET .../results, POST .../cancelPoll a run, read per-case results, or cancel a pending or running run
GET, PUT, DELETE /v1/evals/suites/{suite_id}/scheduleA recurring run every 6, 12, 24, or 168 hours. Pair it with alerts.

A trigger takes target_model, mode (gateway, or external with your own outputs), trials (1 to 5, and a case passes only if every trial passes), metadata (up to 16 string pairs), idempotency_key, and request_overrides for the system prompt and params.

The CI contract. A trigger returns 201 immediately and runs asynchronously. Poll GET /v1/evals/runs/{run_id} until status is completed, failed, or cancelled, then branch on passed. Retrying a trigger with the same idempotency_key returns the original run with 200. Each organization can trigger 10 runs per minute (429 beyond that), and each suite can have 3 runs in progress (409 run_in_flight). See Errors for other codes.

Next steps