Evals API
Run eval suites from CI: sync cases from your repo, trigger runs, poll verdicts
The Evals API runs eval suites from CI: sync cases from your repo, run them against a catalog model or grade your pipeline’s own outputs, and fail the build on a red suite. See the Evals guide for concepts and graders, and running evals from CI for a walkthrough.
Base URL and authentication
Authenticate with the mg_ API key you send /v1/responses with. The organization comes from the key; customer API keys are rejected.
Reference
A trigger takes target_model, mode (gateway, or external with your own outputs), trials (1 to 5, and a case passes only if every trial passes), metadata (up to 16 string pairs), idempotency_key, and request_overrides for the system prompt and params.
The CI contract. A trigger returns 201 immediately and runs asynchronously. Poll GET /v1/evals/runs/{run_id} until status is completed, failed, or cancelled, then branch on passed. Retrying a trigger with the same idempotency_key returns the original run with 200. Each organization can trigger 10 runs per minute (429 beyond that), and each suite can have 3 runs in progress (409 run_in_flight). See Errors for other codes.