Data loss prevention

Detect, redact, or block PII and secrets in prompts before they reach a vendor

Data loss prevention (DLP) scans each prompt against your enabled rules before it reaches the vendor, then records, redacts, or blocks matches, to keep PII and secrets out of vendor traffic or meet a customer DPA. DLP doesn’t scan responses; for leaked credentials in responses, use the prompt injection output check.

Turn on rules

Manage rules under Security → DLP in the dashboard. Every rule, built-in included, starts disabled. Changing rules needs Manage organization settings, and every change is recorded in the audit trail. Built-in rules can be toggled and given a different action, but not edited or deleted.

CategoryEntityDefault action
GlobalCREDIT_CARDRedact
GlobalCRYPTORedact
GlobalDATE_TIMENotify
GlobalEMAIL_ADDRESSRedact
GlobalIBAN_CODERedact
GlobalIP_ADDRESSNotify
GlobalLOCATIONNotify
GlobalPERSONNotify
GlobalPHONE_NUMBERRedact
GlobalURLNotify
USAUS_BANK_NUMBERRedact
USAUS_DRIVER_LICENSERedact
USAUS_ITINRedact
USAUS_PASSPORTRedact
USAUS_SSNRedact

Custom rules

Custom rules match data such as internal IDs, codenames, or your own key formats. Each needs a unique name and at least one matcher, and starts enabled.

FieldLimit
Name80 characters, unique in the organization
PatternA Python regex of up to 500 characters, validated on save
KeywordsUp to 50 case-insensitive strings of up to 100 characters each

Actions

ActionEffect
NotifyRecords the match and sends the prompt unchanged
RedactReplaces each match with a placeholder such as <US_SSN> before the prompt is sent
BlockRejects the request with HTTP 422
{
"error": {
"type": "invalid_request_error",
"message": "Request blocked by DLP policy.",
"code": "blocked_by_dlp_policy"
}
}

On a streaming request, the 422 body has type: "blocked_by_dlp_policy", so match on code. If scanning is unavailable or takes over 250 ms, the request proceeds unscanned with status: "error".

See results on a response

Set include_guardrails_metadata: true in the body, or the include-guardrails-metadata: true header, to get guardrails.dlp on the response, including a blocked 422. It reports entity types and counts, never matched values.

"guardrails": {
"dlp": {
"status": "redacted",
"action": "redact",
"finding_count": 2,
"entity_counts": { "EMAIL_ADDRESS": 1, "US_SSN": 1 },
"action_counts": { "REDACT": 2 },
"rule_ids": ["6f1c2a8e-4b7d-4e0a-9c3f-2d8b1a7e5f60"]
}
}

status is one of clean, flagged, redacted, blocked, skipped (with a skipped_reason such as no_enabled_rules), or error. action is none, notify, redact, or block.

Test rules

Security → DLP tester runs up to 50,000 characters against your enabled rules and lists each hit’s entity, action, matched text, and offsets. It uses regex approximations, so it never reports PERSON, LOCATION, US_PASSPORT, US_DRIVER_LICENSE, or US_ITIN, which live traffic does detect. Confirm those with a real request and guardrails metadata.

Alerts

Every match appears under Security → Alerts with the entity, action, model, and customer ID when present, alongside prompt injection detections.

Override per project or customer

Projects and customers can override the action per entity type.

Next steps