Telemetry export

Send your Gateway request telemetry to your own observability backend

Telemetry export sends one span per Gateway request to your own observability backend, so LLM traffic sits beside the rest of your system with no code changes. Spans go out as OTLP over HTTP/protobuf, following the OpenTelemetry GenAI semantic conventions.

Set up a destination

Open Settings → Telemetry export and click Add destination. This needs Manage telemetry export (Admin and Developer roles).

DestinationTarget fieldCredentialTraces land in
DatadogDatadog site: datadoghq.com, us3.datadoghq.com, us5.datadoghq.com, datadoghq.eu, ap1.datadoghq.com, or ap2.datadoghq.comAPI keyLLM Observability
Grafana CloudZone, such as prod-us-east-0 from otlp-gateway-prod-us-east-0.grafana.netInstance ID and Access policy tokenTempo
Custom (OTLP)Endpoint URL, the base URL only. Gateway appends /v1/traces and /v1/metrics.Bearer token, sent as Authorization: BearerYour endpoint

Custom (OTLP) needs Enterprise and accepts any OTLP/HTTP receiver: an OpenTelemetry Collector, Honeycomb, New Relic, SigNoz, Langfuse, or self-hosted Tempo or Jaeger. It must use https and can’t redirect or resolve to a private, loopback, link-local, reserved, or cloud metadata address.

ML app (Service name for Grafana and custom) sets service.name, defaulting to the destination name. Each enabled destination gets all your organization’s traffic, so a second one copies rather than splits it.

Prompts and completions

Exclude prompts and completions is on by default. It strips gen_ai.input.messages, gen_ai.output.messages, gen_ai.tool.definitions, and error messages (which can quote the prompt); models, tokens, latency, and cost still arrive. Turning it off sends content to that destination, independent of Gateway’s own payload logging. It’s re-read before each batch, so re-enabling it stops content immediately, even for queued requests.

Metrics

Send metrics (off by default) adds:

MetricType
gen_ai.client.operation.durationHistogram, seconds
gen_ai.client.token.usageHistogram, by gen_ai.token.type
gen_ai.client.time_to_first_tokenHistogram, seconds
merge.cost.usdSum
Leave metrics off for Grafana Cloud

Metrics use delta temporality, which Datadog requires and Grafana Cloud’s OTLP gateway rejects with invalid temporality and type combination. The dashboard still allows the toggle, but points are rejected (traces still reach Tempo). Derive span metrics with Tempo’s metrics generator instead.

Spans carry the same numbers and Datadog bills custom metrics separately, so enable metrics for long retention, cheap wide-range queries, or alerts. In Datadog, attribute keys keep their dots (group by {merge.cost.basis}), and histogram percentile queries need percentiles enabled on the metric first.

Delivery

Gateway batches spans per destination and retries transient failures with backoff; a rejected credential or malformed endpoint isn’t retried. Repeated failures pause that destination for a while before Gateway retries. Export runs outside the request path, so a failing destination never affects your LLM calls. Spans usually arrive within a few minutes, plus a few more in Datadog LLM Observability.

Reference

AttributeMeaning
gen_ai.operation.namechat, embeddings, image_generation, speech_synthesis, or audio_transcription
gen_ai.request.model, gen_ai.response.modelRequested and served model
gen_ai.provider.nameSemconv provider name, such as aws.bedrock
gen_ai.usage.input_tokens, gen_ai.usage.output_tokensToken counts
gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_creation.input_tokensPrompt cache activity
gen_ai.conversation.idSession ID, when the request carried one
gen_ai.input.messages, gen_ai.output.messages, gen_ai.tool.definitionsContent, only when prompts and completions aren’t excluded
merge.request_id, merge.project_id, merge.customer_idGateway request, project, and customer IDs
merge.vendor, merge.was_fallback, merge.fallback_reasonServing vendor and whether a fallback served the request
merge.cost.usd, merge.cost.basisCost and what it measures
error.typeError code, on failed requests

merge.cost.basis says whose charge merge.cost.usd is:

Basismerge.cost.usd is
merge_billedWhat you pay Merge for the request
provider_list_priceThe vendor’s list price for BYOK traffic, which you pay your vendor
unavailableOmitted, because Gateway couldn’t price the request

Sum both priced bases for total LLM spend. A merge_billed 0 means the request used no tokens.

Next steps