Evals API
The Evals API runs suites and migrations from automation, such as a CI job that blocks a deploy on a failing suite. Schemas are in the Evals API and Migrations API reference. Authenticate with your organization-level mg_ key (the one you use for /v1/responses); customer API keys are rejected.
Gate CI on a suite
Start a run, poll until status leaves pending or running, then branch on passed (there’s no blocking mode).
A run ends completed, failed (couldn’t execute), or cancelled; passed is null unless it completed. Completed runs also carry pass_rate, its 95% interval (ci_low, ci_high), total_cost_usd, and trace_id.
Trigger fields
Dataset as code
Keep cases in your repo so a prompt change ships with its cases. PUT /v1/evals/suites creates (201) or updates (200) a suite by name. PUT /v1/evals/suites/{suite_id}/cases syncs cases by external_id: it updates matches, creates new ones, and deletes synced cases missing from the payload unless you send "prune": false.
Syncs never touch dashboard-authored cases (no external_id), and overwrite dashboard edits to a Synced case.
External runs
When your pipeline runs the model itself (a local checkpoint or your full agent loop), send "mode": "external" with one outputs entry per case, matched by external_id or case name and carrying output_text, output_json, or tool_calls.
target_model is a free-text label, trials are always 1, LLM judge graders still bill, and an enabled case with no output fails.
The gate uses the latest completed run of either mode whose target_model matches the candidate. An external run labeled with the candidate’s exact model ID sets the verdict without the candidate running through Gateway, so pick another label unless you intend that.
Schedules
GET, PUT, and DELETE on /v1/evals/suites/{suite_id}/schedule manage the schedule. PUT takes target_model, interval_hours (6, 12, 24, or 168), trials, and is_enabled, and restarts the clock.
Migrations
POST /v1/migrations creates a migration draft from baseline_model, candidate_model, and experiment_suite_id, with optional shadow_sample_rate (above 0 up to 1, default 1) and shadow_daily_budget_usd. PATCH takes exactly one change: a status (such as shadowing or paused), an experiment_suite_id (null unlinks), or a shadow_sample_rate.
To cut over, POST /v1/migrations/{migration_id}/complete with up to 50 policy_ids from GET /v1/migrations/{migration_id}/affected-policies (an empty list changes no traffic). Revert with PATCH {"status": "reverted"}.
Any organization-level Gateway key can complete a migration and change production routing, so protect automation keys like a deploy credential.
Webhook alerts
Add one webhook per organization under Evals → Alerts. Each new alert POSTs JSON with kind (run_failed, run_error, or regression), title, suite_id, run_id, migration_id, and detail. X-Merge-Signature: sha256=<hex> is an HMAC-SHA256 of the raw body under the signing secret, shown once when you create or rotate the webhook.
Any 2xx acknowledges; failures retry four times over about 75 minutes.
Errors
Eval errors return {"detail": {"code": "...", "message": "..."}}. Most migration errors return detail as a plain string, so branch on status.