Skip to main content
Version: 1.7.0

Observability

Runs unlicensed

Runs unlicensed, in full. The tamper-evident signed hash chain, SIEM export and compliance reports are available in the Development posture, with no reduced retention. Earlier releases kept only a local unsigned log here; 1.7.0 removed that split.

DVARA provides structured JSON logging, Prometheus metrics, token usage metering, and request tracing out of the box.

X-Trace-ID Propagation

Every response includes an X-Trace-ID header for request correlation. This applies to every plane the gateway serves on port 8080 — LLM, MCP and A2A.

Behavior:

  • If the incoming request includes an X-Trace-ID header, the same value is echoed back.
  • When OpenTelemetry tracing is active, the OTel 32-character hex trace ID is used if no client header is present.
  • Otherwise, the gateway generates a new random 32-character hex ID.
  • The trace ID is embedded in every error response body as error.trace_id.
  • The trace ID is added to SLF4J MDC as trace_id for inclusion in all structured log lines.
# With custom trace ID
curl -i -H "Authorization: Bearer $DVARA_API_KEY" -H "X-Trace-ID: my-custom-trace-001" \
http://localhost:8080/v1/models
# → X-Trace-ID: my-custom-trace-001

# Without trace ID (auto-generated)
curl -i -H "Authorization: Bearer $DVARA_API_KEY" http://localhost:8080/v1/models
# → X-Trace-ID: a6783439db1f46a6bfed511a0011e955

X-Session-Id Header

Both the LLM gateway and MCP Gateway accept an optional X-Session-Id header for agent session correlation. When present, the session ID is:

  • Stored as a servlet request attribute (sessionId)
  • Added to SLF4J MDC as session_id for structured log correlation
  • Attached as a high-cardinality attribute on OpenTelemetry spans

This allows traces from multiple LLM turns and MCP tool calls within a single agent session to be correlated by session ID.

# LLM request with session ID
curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $DVARA_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Session-Id: agent-session-42" \
-d '{"model":"gpt-4o","messages":[{"role":"user","content":"Hello"}]}'

# MCP request with same session ID
curl http://localhost:8080/mcp/filesystem/tools/call \
-H "Content-Type: application/json" \
-H "Authorization: Bearer gw_mykey" \
-H "X-Session-Id: agent-session-42" \
-d '{"name":"read_file","arguments":{"path":"/data/file.txt"}}'

Structured JSON Logging

DVARA emits structured JSON logs by default. Every log line is a JSON object, ready to ingest into ELK, Loki, Datadog, Splunk, or any other JSON-aware log pipeline.

Configuration

  • JSON mode (default) — Every log line is a JSON object with @timestamp, level, message, logger_name, and MDC fields.
  • Plain-text mode — Human-readable output for local development. Activate with the log-plain profile (SPRING_PROFILES_ACTIVE=log-plain).

Log fields

Every access log line and every gateway log line within the scope of a request carry these fields:

FieldDescription
trace_idRequest correlation ID (see above)
session_idAgent session ID (from X-Session-Id header, if present)
workspace_idWorkspace identifier
modelRequested model name
providerSelected provider
methodHTTP method
pathRequest URI
statusHTTP response status
latency_msRequest duration in milliseconds
api_keyAPI key (masked to first 8 chars)
cache_statusHIT or MISS
streamWhether request was streaming
tokens_promptPrompt token count
tokens_completionCompletion token count
tokens_totalTotal token count
error_codeGateway error code (if any)
priority_tierPriority admission tier the request was admitted under (premium / standard / bulk), present only when priority admission control is enabled

Access Log Example

Every request produces a single structured access log entry:

{
"@timestamp": "2026-02-25T10:23:45.123Z",
"level": "INFO",
"message": "Request completed",
"logger_name": "dvara.access",
"trace_id": "a6783439db1f46a6bfed511a0011e955",
"workspace_id": "acme-corp",
"model": "gpt-4o",
"provider": "openai",
"method": "POST",
"path": "/v1/chat/completions",
"status": "200",
"latency_ms": "342",
"api_key": "sk-prod-1...",
"cache_status": "MISS",
"tokens_prompt": "150",
"tokens_completion": "85",
"tokens_total": "235",
"service": "dvara-gateway"
}

Prometheus Metrics

DVARA exposes a Prometheus-format metrics endpoint out of the box.

Scrape Endpoint

GET /actuator/prometheus
Authorization: Bearer $DVARA_ACTUATOR_METRICS_API_KEY

The endpoint is authenticated — configure your Prometheus job's bearer_token_file to point at the DVARA_ACTUATOR_METRICS_API_KEY value. The metrics secret is intentionally distinct from DVARA_ACTUATOR_API_KEY (which guards /actuator/gateway-status) so a leaked scrape token can't unlock the license envelope. A worked Prometheus scrape config (bearer_token_file placement, metrics_path: /actuator/prometheus) is not included here; for now, configure Prometheus per its standard documentation pointing at the gateway's /actuator/prometheus endpoint with the metrics-key bearer.

Core Metrics

MetricTypeLabelsDescription
gateway_requests_totalCounterworkspace, model, provider, status, regionTotal gateway requests
gateway_latency_secondsHistogramworkspace, model, provider, status, regionRequest latency (P50/P95/P99)
gateway_tokens_totalCounterworkspace, model, directionToken usage (direction=input/output)
gateway_provider_errors_totalCounterprovider, error_codeProvider errors
gateway_retries_totalCounterproviderRetry attempts
gateway_fallbacks_totalCounterfrom_provider, to_providerFallback activations

Policy & Routing Metrics

MetricTypeLabelsDescription
gateway_policy_shadow_divergence_totalCounterpolicy_id, rule_id, divergence_typeShadow policy divergence events
gateway_canary_requests_totalCounterroute_id, variant, modelCanary A/B test requests
gateway_shadow_requests_totalCounterroute_id, primary_provider, shadow_providerShadow traffic routing events
gateway_priority_requests_totalCounterworkspace, tierPriority-routed requests
gateway_priority_throttled_totalCounterworkspace, tierPriority admission rejections

FinOps Metrics

MetricTypeLabelsDescription
gateway_cost_dollars_totalCounterworkspace, model, providerCumulative cost in USD
gateway_budget_blocked_totalCounterworkspace, budget_idHard budget cap rejections
gateway_budget_soft_alert_totalCounterworkspace, budget_idSoft budget cap alerts
gateway_budget_warning_totalCounterworkspace, budget_idBudget warning policy triggers
gateway_model_downgrades_totalCounterworkspace, original_model, downgraded_modelAutomatic model downgrades
gateway_cost_anomaly_totalCounterworkspace, modelCost anomaly detections

Guardrail & Security Metrics

MetricTypeLabelsDescription
gateway_guardrail_blocked_totalCounterworkspace, categoryGuardrail BLOCK actions
gateway_guardrail_flagged_totalCounterworkspace, categoryGuardrail FLAG actions
gateway_schema_validations_totalCounterworkspace, model, resultOutput schema validations
gateway_schema_retries_totalCounterworkspace, modelSchema validation retries
gateway_context_window_warnings_totalCounterworkspace, modelContext window threshold warnings
gateway_context_window_pruned_totalCounterworkspace, model, strategyContext window pruning events
gateway_mcp_injection_detections_totalCounterworkspace, server_id, actionMCP injection detections
gateway_ml_guardrail_totalCounterprovider, category, actionML-classifier guardrail decisions (Lakera, ShieldGemini, Bedrock Guardrails, Aporia, onnx-injection)
gateway_plugin_guardrail_totalCounterplugin, category, actionExternal guardrail plugin decisions
gateway_grounding_check_totalCountergrounded, actionEmbedding-based grounding check outcomes

Intelligent Routing & Config Metrics

MetricTypeLabelsDescription
gateway_intelligent_routing_totalCountercomplexity, selected_modelIntelligent routing selections by complexity tier
gateway_config_import_totalCountermode, dry_runGitOps config import operations

Flightdeck-only metrics

These counters live in the flightdeck app's MeterRegistry (not in gateway-server's), so they're only scrapeable from flightdeck:8090/actuator/prometheus with the flightdeck's own DVARA_ACTUATOR_METRICS_API_KEY. Configure a second Prometheus job pointed at the flightdeck pod if you want email-pipeline health on your dashboards.

MetricTypeLabelsDescription
dvara_emails_sent_totalCountertemplate, transport, result, deliveredEmail send attempts. resultSUCCESS / TRANSIENT / PERMANENT / MAX_ATTEMPTS_EXCEEDED. transportlog / smtp / resend. deliveredtrue / false — whether the message left the process. Query real deliveries with delivered="true"; result="SUCCESS" alone counts console prints on the default log transport.
dvara_emails_retried_totalCountertemplate, attemptRetries scheduled by the durability layer (counts retry attempt N, not the original send).

MCP Gateway Metrics

MetricTypeLabelsDescription
mcp_tool_calls_totalCounterworkspace, server_id, tool_name, statusMCP tool call count
mcp_tool_call_latency_secondsHistogramserver_id, tool_nameMCP tool call latency
mcp_approval_requests_totalCounterworkspace, server_id, tool_nameApproval gate requests
mcp_approval_granted_totalCounterworkspaceApprovals granted
mcp_approval_denied_totalCounterworkspaceApprovals denied
mcp_approval_timeout_totalCounterworkspaceApproval timeouts
mcp_agent_loop_detected_totalCounterworkspace, loop_typeAgent loop detection fires
mcp_agent_sessions_killed_totalCounterworkspaceAgent sessions killed

Data-Plane Channel & Durability Metrics

These cover a gateway pod that receives configuration as a signed bundle and spools its audit records locally. They are the ones worth alerting on, because each failure they describe is silent — the pod keeps serving traffic while drifting out of date or falling behind.

MetricTypeWhat it tells you
gateway_audit_chain_gaps_totalCounterBreaks detected in the audit hash chain. Any non-zero value deserves investigation — on a hosted install a break at the retention horizon is the retention sweep, but anywhere else it is not expected.
gateway_config_bundle_age_secondsGaugeHow old the configuration this pod is serving is. Rising steadily means it has stopped receiving updates.
gateway_config_bundle_seconds_since_contactGaugeHow long since the pod last reached the control plane.
gateway_config_bundle_ever_contactedGauge0 means this pod has never reached the control plane — usually a client-certificate problem, and distinct from "contacted and now stale".
gateway_deny_list_seconds_since_contactGaugeHow long since the pod refreshed the revocation list. A revocation issued now would not reach it.
gateway_deny_list_ever_contactedGauge0 means the revocation channel was never reached at all.
gateway_spool_drained_totalCounterAudit records shipped from the local spool to Flightdeck.
gateway_spool_drain_failures_totalCounterFailed drain attempts. Sustained non-zero means audit records are accumulating on disk.
gateway_scheduler_last_success_timestamp_secondsGaugeWhen each scheduled job last succeeded. A stalled timestamp is a job that stopped without failing loudly.
gateway_config_poller_errors_totalCounterFailed config polls.
gateway_config_poller_areas_refreshed_totalCounterConfig areas refreshed after a change was detected.
gateway_config_change_received_totalCounterConfig changes the pod acted on. mcp_ and a2a_ variants exist for those planes.
gateway_hazelcast_runningGaugeWhether the embedded cache cluster is up — it backs rate limiting and the config cache.
gateway_license_days_remainingGaugeDays until the licence expires.
gateway_siem_export_totalCounterAudit events exported to your SIEM.
gateway_quota_counter_seed_totalCounterHow the fleet-wide quota counter was seeded on start — a rising zero source across the fleet means quotas are being enforced against no history.
gateway_quota_seed_checkpoint_age_secondsGaugeHow stale the seed was. Exposure to over-use is roughly this multiplied by call rate.
gateway_quota_checkpoint_cellsGaugeCells in the last checkpoint written. Flat at 0 means no checkpoint is being taken.
dvara_ingest_clock_skew_clamped_totalCounterSpooled records whose timestamps were clamped because a pod's clock disagreed with the control plane.

Configuration

Metrics are enabled by default in application.yml. The exact exposure.include list differs per app:

# gateway-server (port 8080)
management:
endpoints:
web:
exposure:
include: health,prometheus,gateway-status,info
prometheus:
metrics:
export:
enabled: true
AppDefault include
gateway-server (port 8080)health,prometheus,gateway-status,info
flightdeck (port 8090)health,prometheus,info

Dropping gateway-status from the gateway-server list breaks the Flightdeck License page (it reads license metadata from /actuator/gateway-status); dropping info breaks anonymous build-info polling by uptime monitors.

Grafana Dashboard

Example PromQL queries:

# Request rate by provider
rate(gateway_requests_total[5m])

# P95 latency by model
histogram_quantile(0.95, rate(gateway_latency_seconds_bucket[5m]))

# Token throughput by workspace
rate(gateway_tokens_total[5m])

# Error rate by provider
rate(gateway_provider_errors_total[5m])

Token Usage Metering

Every non-streaming chat request records token usage to the DVARA token ledger, backed by PostgreSQL. Records are on the Console's Token Usage page, described below.

Query Endpoints

Browse the token ledger on the Console's Token Usage page (/token-usage), filtering by workspace, API key, model and date range. The page shows a summary — input, output and total tokens, request count, average tokens per request — above the records themselves, and marks rows whose counts were estimated rather than reported by the provider.

For a billing-grade monthly figure, use the export on that page: it separates served volume (exact, estimated and cache-hit) from upstream cost, which excludes cache hits.

Record Fields

FieldTypeDescription
idstringUnique record ID (UUID)
tenantIdstringWorkspace identifier
apiKeystringAPI key used
modelstringModel requested
providerstringProvider that served the request
inputTokensintPrompt tokens
outputTokensintCompletion tokens
totalTokensintTotal tokens
estimatedbooleanWhether counts are estimated
timestampISO 8601When the request was made

Summary Response

{
"tenantId": "acme-corp",
"model": "gpt-4o",
"totalInputTokens": 15000,
"totalOutputTokens": 8500,
"totalTokens": 23500,
"requestCount": 42
}

Pre-Built Grafana Dashboards

Five production-ready Grafana dashboards ship in the dvarahq/dvara-examples repo under grafana/dashboards/ — not in the gateway distribution itself. Clone or download that repo to use them; every path below is relative to its root.

DashboardFileDescription
Gateway Overviewdvara-overview.jsonRequest volume, latency (P50/P95/P99), error rates, provider health, token usage, cost/hour
FinOps & Budgetdvara-finops.jsonCost by workspace/model/provider, budget enforcement, model downgrades, anomalies
MCP Gateway & Agenticdvara-mcp.jsonTool calls, agent sessions, loop detection, approval gates, injection detection
Policy, Routing & Fleetdvara-policy-routing.jsonShadow policy divergence, canary testing, priority routing, config sync, fleet health
Infrastructuredvara-infrastructure.jsonJVM, connection pool, distributed cache, PostgreSQL — gateway-host process health

One-Command Setup

From inside the cloned dvara-examples repo:

docker compose -f docker-compose.yml -f grafana/docker-compose.monitoring.yml up

This starts Prometheus (port 9090) and Grafana (port 3000, admin/dvara) with dashboards auto-provisioned.

Manual Import

for f in grafana/dashboards/*.json; do
curl -X POST http://admin:dvara@localhost:3000/api/dashboards/db \
-H "Content-Type: application/json" \
-d "{\"dashboard\": $(cat "$f"), \"overwrite\": true}"
done

Alerting Rules

grafana/alerts/dvara-alerts.yml in dvarahq/dvara-examples defines 12 Prometheus alerting rules:

AlertSeverityTrigger
DvaraHighErrorRatecriticalError rate > 5% for 5 min
DvaraHighP95LatencywarningP95 > 5s for 5 min
DvaraProviderErrorSpikewarningProvider errors > 1/sec for 3 min
DvaraCircuitBreakerOpencriticalErrors with zero successes for 2 min
DvaraBudgetHardLimitcriticalHard budget cap hit
DvaraBudgetSoftLimitwarningSoft limit breached
DvaraCostAnomalywarningCost exceeds baseline
DvaraGuardrailBlockswarningGuardrail blocks > 0.1/sec
DvaraAgentLoopDetectedwarningAgent loop detected
DvaraApprovalTimeoutswarningApproval gate timeouts
DvaraMcpToolErrorRatewarningMCP error rate > 10%
DvaraInjectionDetectedcriticalPrompt injection detected

Alert names are Title-cased DVARA*, not all-caps DVARA* — match them exactly when wiring into PagerDuty / OpsGenie routing rules or runbook entries.

Datadog Integration

The Datadog assets — Agent config and pre-built monitors — also ship in the dvarahq/dvara-examples repo, under datadog/. Clone that repo to use them.

OpenMetrics Scraping

Copy the Datadog Agent config to scrape all DVARA Prometheus metrics:

cp datadog/conf.d/dvara.yaml /etc/datadog-agent/conf.d/openmetrics.d/dvara.yaml
sudo systemctl restart datadog-agent

OTLP Traces to Datadog

Configure the Datadog Agent as an OTLP collector, then point DVARA to it:

OTEL_EXPORTER_OTLP_ENDPOINT=http://datadog-agent:4318/v1/traces

Pre-built Datadog monitors are provided in datadog/monitors.yaml.

Health Endpoints

The gateway uses a two-key model for authenticated actuator endpoints, with the probe paths left anonymous so container orchestrators don't need credentials. The worked Prometheus bearer_token_file scrape config is not included here.

EndpointAuthDescription
GET /actuator/healthAnonymous (permitAll)Liveness summary. Returns {"status":"UP|DOWN"} only — detail is gated by management.endpoint.health.show-details=when-authorized.
GET /actuator/health/livenessAnonymous (k8s probe)Liveness probe target.
GET /actuator/health/readinessAnonymous (k8s probe)Readiness probe target — fails closed when the database, license, or any other registered indicator reports DOWN.
GET /actuator/infoAnonymousBuild info (app.*, build.*).
GET /actuator/gateway-statusAuthorization: Bearer $DVARA_ACTUATOR_API_KEYRich gateway status: mode, version, providers, routes, rate limits, license, warnings, uptime. Powers the DVARA Flightdeck License page.
GET /actuator/prometheusAuthorization: Bearer $DVARA_ACTUATOR_METRICS_API_KEYPrometheus scrape endpoint. The metrics secret is intentionally distinct from DVARA_ACTUATOR_API_KEY — by the principle of least privilege, a leaked scrape token must not unlock the gateway-status surface.
/actuator/gateway-status exposes metadata only

Even with DVARA_ACTUATOR_API_KEY, the gateway-status endpoint returns license metadata (licensee name, expiry, ID, runtime status) — never the raw DVARA_LICENSE_KEY envelope. Once validated at startup the envelope stays in process memory; no API surfaces it. Provider API keys, vault credentials, and the audit HMAC secret are likewise never returned by this endpoint — only operator-facing metadata.

| GET /v1/models | Workspace API key | Lists all registered providers and models with capabilities. |

/actuator/env, /heapdump, /threaddump, /beans, /mappings, /configprops, /loggers, /scheduledtasks, /caches, /sessions, and /quartz are excluded from the registry. The status depends on who is asking: 401 anonymously — the security chain rejects before the registry is consulted — and 404 once authenticated. Both are correct, and neither tells an unauthenticated scanner whether the endpoint exists.

OpenTelemetry Distributed Tracing

DVARA includes OpenTelemetry distributed tracing out of the box. Traces are automatically created for every request and provider call, with W3C traceparent headers propagated to upstream LLM providers.

How It Works

Both the DVARA LLM Gateway and the DVARA MCP Gateway have full OTLP tracing support:

  1. Server spans — one per incoming HTTP request
  2. LLM provider spans — one per provider call (chat, streamChat, embed), enriched with token usage and session ID
  3. MCP filter spans — one per stage of the MCP filter pipeline (server registry lookup, policy evaluation, PII scanning, upstream server call)
  4. Client spans — one per outbound HTTP call, with W3C traceparent headers propagated to upstream LLM and MCP providers

LLM Span Hierarchy

HTTP POST /v1/chat/completions (server span)
└── gateway.provider.chat (gateway observation)
├── low-card: provider, model
├── high-card: input_tokens, output_tokens, session_id
└── HTTP POST https://api.openai.com (client span)

MCP Span Hierarchy

HTTP POST /mcp/{serverId}/tools/call (server span)
└── dvara.mcp-gateway.request (parent MCP observation)
├── low-card: server_id, operation, tool_name
├── high-card: session_id, latency_ms, response_bytes, pii_in_response

├── mcp.filter.registry (low-card: server_id)
├── mcp.filter.policy (low-card: decision)
├── mcp.filter.pii_args (low-card: action)
├── mcp.server.call (upstream HTTP call)
│ ├── low-card: server_id, operation, tool_name
│ ├── high-card: http_status, latency_ms, response_bytes
│ ├── event: mcp_request_sent
│ └── event: mcp_response_received
└── mcp.filter.pii_response (low-card: action, conditional)

LLM Span Names and Attributes

Observation NameOperationLow-CardinalityHigh-Cardinality
gateway.provider.chatNon-streaming chatprovider, modelinput_tokens, output_tokens, session_id
gateway.provider.streamStreaming chatprovider, modelsession_id
gateway.provider.embedEmbeddingsprovider, model

These same spans also carry OpenTelemetry GenAI semantic-convention attributes (gen_ai.*), so LLM-aware backends (Langfuse, Arize Phoenix, …) render model and token usage natively — see GenAI semantic-convention attributes below.

MCP Span Names and Attributes

Observation NameOperationLow-CardinalityHigh-Cardinality
dvara.mcp-gateway.requestParent MCP requestserver_id, operation, tool_namesession_id, latency_ms, response_bytes, pii_in_response
mcp.filter.registryServer registry lookupserver_id
mcp.filter.policyPolicy evaluationdecision
mcp.filter.pii_argsPII scan on request argsaction
mcp.filter.pii_responsePII scan on responseaction
mcp.server.callUpstream HTTP callserver_id, operation, tool_namehttp_status, latency_ms, response_bytes

Configuration

management:
tracing:
sampling:
probability: ${TRACING_SAMPLING_PROBABILITY:1.0} # 0.0–1.0, default: sample all
opentelemetry:
tracing:
export:
otlp:
endpoint: ${OTEL_EXPORTER_OTLP_ENDPOINT:http://localhost:4318}
Environment VariableDefaultDescription
TRACING_SAMPLING_PROBABILITY1.0Fraction of traces to sample (0.0 = none, 1.0 = all)
OTEL_EXPORTER_OTLP_ENDPOINThttp://localhost:4318OTLP base URL for trace export
The OTLP endpoint is a base URL — no /v1/traces suffix

OTEL_EXPORTER_OTLP_ENDPOINT is the collector's base URL (e.g. http://otel-collector:4318); the OpenTelemetry SDK appends the /v1/traces signal path itself. Do not include /v1/traces — that produces POST …/v1/traces/v1/traces and the collector rejects it (405). The Spring property is management.opentelemetry.tracing.export.otlp.endpoint.

X-Trace-ID Integration

When OpenTelemetry tracing is active, the X-Trace-ID response header uses the OTel 32-character hex trace ID instead of a random UUID. Client-supplied X-Trace-ID headers still take precedence.

ScenarioX-Trace-ID Value
Client sends X-Trace-ID headerClient's value (echoed back)
OTel tracing active, no client headerOTel trace ID (32-char hex)
No tracing, no client headerRandom UUID hex (32-char)

Viewing Traces

Start a local Jaeger instance for trace visualization:

# Start Jaeger with OTLP collector
docker run -d --name jaeger -p 16686:16686 -p 4318:4318 jaegertracing/all-in-one:latest

# Start the gateway with traces auto-exported to localhost:4318 (base URL — no /v1/traces)
docker run -d --name dvara-gateway \
-p 8080:8080 \
-e MOCK_PROVIDER_ENABLED=true \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://host.docker.internal:4318 \
ghcr.io/dvarahq/dvara-gateway:1.7.0

# Send a request
curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $DVARA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"mock/test","messages":[{"role":"user","content":"Hello"}]}'

# View traces at http://localhost:16686

GenAI semantic-convention attributes (gen_ai.*)

DVARA emits OpenTelemetry GenAI semantic-convention attributes on the LLM provider spans, alongside the native provider / model / token attributes. GenAI-aware backends (Langfuse, Arize Phoenix, Weave, …) then parse model, request parameters, and token usage natively — no vendor-specific exporter, just point OTLP at them.

Spangen_ai.* attributes
gateway.provider.chatgen_ai.system, gen_ai.request.model, gen_ai.operation.name (chat), gen_ai.request.max_tokens, gen_ai.request.temperature, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.model, gen_ai.response.id, gen_ai.response.finish_reasons
gateway.provider.streamgen_ai.system, gen_ai.request.model, gen_ai.operation.name (chat), gen_ai.request.max_tokens, gen_ai.request.temperature
gateway.provider.embedgen_ai.system, gen_ai.request.model, gen_ai.operation.name (embeddings)

gen_ai.system maps the DVARA provider to its semantic-convention value where one is enumerated — openai, anthropic, aws.bedrock, az.ai.openai, gcp.gemini, cohere, mistral_ai, deepseek, groq, xai — and falls back to the lower-cased provider name for the OpenAI-compatible long-tail (Qwen, Moonshot, ChatGLM, Ollama, …).

To suppress the gen_ai.* attributes (e.g. if the extra span cardinality is a concern), set management.tracing.genai-semconv.enabled: false (default true). The native provider / model / token attributes remain regardless.

Sending traces to Arize Phoenix

Phoenix ingests OTLP directly — run it locally, point DVARA at it, and the gen_ai.* attributes appear natively.

# Start Phoenix (UI + OTLP receiver both on 6006)
docker run -d --name phoenix -p 6006:6006 arizephoenix/phoenix:latest

# Point the gateway at Phoenix — base URL only (the SDK appends /v1/traces)
docker run -d --name dvara-gateway \
-p 8080:8080 \
-e MOCK_PROVIDER_ENABLED=true \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://host.docker.internal:6006 \
ghcr.io/dvarahq/dvara-gateway:1.7.0

# Send a request, then open the Phoenix UI
curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $DVARA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"mock/test","messages":[{"role":"user","content":"Hello"}]}'

# View traces at http://localhost:6006 — the gateway.provider.chat span shows
# gen_ai.system / gen_ai.request.model / gen_ai.usage.* natively.

For a hosted Phoenix, set OTEL_EXPORTER_OTLP_ENDPOINT to the instance's base URL.

Sending traces to Langfuse

Langfuse exposes an OTLP endpoint at /api/public/otel; traces authenticate with a project key pair sent as an HTTP Basic Authorization header via the standard OpenTelemetry headers variable.

docker run -d --name dvara-gateway \
-p 8080:8080 \
-e OTEL_EXPORTER_OTLP_ENDPOINT=https://cloud.langfuse.com/api/public/otel \
-e OTEL_EXPORTER_OTLP_HEADERS="Authorization=Basic <base64(public-key:secret-key)>" \
ghcr.io/dvarahq/dvara-gateway:1.7.0
  • Langfuse Cloudhttps://cloud.langfuse.com/api/public/otel (EU) or https://us.cloud.langfuse.com/api/public/otel (US); the header value is Basic + base64 of <public-key>:<secret-key>.
  • Self-hosted — use your instance's /api/public/otel base URL.

The same gen_ai.* attributes surface in Langfuse's model / usage views.

Verification

The Phoenix flow above is verified end-to-end. The Langfuse endpoint and header values follow Langfuse's published OTLP integration — confirm the header binding in your own deployment.

Log Correlation

When tracing is active, OTel trace and span IDs are automatically added to every log line in the request scope:

{
"trace_id": "abcdef1234567890abcdef1234567890",
"traceId": "abcdef1234567890abcdef1234567890",
"spanId": "1234567890abcdef",
"span_id": "1234567890abcdef",
"message": "Request completed",
...
}

Disabling Tracing

Set the sampling probability to 0.0 to disable trace collection while keeping the infrastructure in place:

TRACING_SAMPLING_PROBABILITY=0.0 docker run -d ghcr.io/dvarahq/dvara-gateway:1.7.0

Audit Event Stream

Every API request through /v1/* generates audit events that are persisted and queryable.

Audit event catalogue

The events below are the ones customers and auditors ask about most often: who was allowed to do what, when personal data was seen, and who changed the rules.

This is a selection, not the complete set

DVARA writes more event types than are listed here — around 130 in total, including administrative and internal-operations events. SIEM export carries every one of them, not only these. This page lists the subset that answers the questions usually put to it; if you need an event that is not here, it is still being recorded and still exported.

Every event shares the same envelope — eventId, timestamp, workspaceId, eventType, payload — and is HMAC-signed and hash-chained.

Someone was allowed, or refused, to act.

EventWritten when
MCP_APPROVAL_REQUESTEDA tool call hit an approval gate and is waiting for a human
MCP_APPROVAL_GRANTEDA reviewer approved it
MCP_APPROVAL_DENIEDA reviewer refused it
MCP_APPROVAL_TIMEOUTNobody answered in time, so the default action applied
MCP_APPROVAL_RESOLVEDThe pending request reached a final state
A2A_APPROVAL_RESOLVEDThe same, for an agent-to-agent hop
POLICY_DENIEDA request was refused by a policy rule
A2A_POLICY_DENIEDAn agent-to-agent hop was refused by a policy rule
GUARDRAIL_BLOCKEDA request or reply was refused by a guardrail
GUARDRAIL_FLAGGED / GUARDRAIL_DETECTEDA guardrail matched but did not block
SESSION_KILLEDAn agent session was terminated
AGENT_LOOP_DETECTEDAn agent was caught repeating itself

Personal data was seen.

EventWritten when
PII_DETECTEDPersonal data was found in a request
PII_REDACTEDIt was replaced with a placeholder before going to the provider
PII_OUTPUT_LEAKPersonal data was found in the reply coming back
MCP_PII_DETECTED / MCP_PII_REDACTED / MCP_PII_OUTPUT_LEAKThe same three, on a tool call rather than a model call

None of these carries the personal data itself — the payload records the entity types and counts, never the values.

Somebody changed the rules.

EventWritten when
POLICY_CREATED / POLICY_UPDATED / POLICY_DELETEDA policy was written or removed
POLICY_STATUS_CHANGED / POLICY_PROMOTED / POLICY_ROLLED_BACKA policy was activated, promoted from shadow, or reverted
WORKSPACE_PII_CONFIG_UPDATEDA workspace's PII settings changed
WORKSPACE_GUARDRAIL_CONFIG_UPDATEDA workspace's guardrail settings changed
GUARDRAIL_PLUGIN_CREATED / _UPDATED / _ROTATED / _DELETEDAn external guardrail plugin was registered, changed, or had its secret rotated

Somebody gained or lost access.

EventWritten when
USER_INVITED / USER_DELETEDA person was given or removed access
API_KEY_CREATED / API_KEY_REVOKEDA key was minted or revoked
WORKSPACE_API_KEYS_REVOKEDEvery key in a workspace was revoked at once
WORKSPACE_CREATED / WORKSPACE_UPDATED / WORKSPACE_DELETEDA workspace was created, changed, or removed
WORKSPACE_SWITCHEDA user changed which workspace they were viewing
WORKSPACE_SWITCH_REFUSEDA user tried to reach a workspace outside their team and was refused
PROVIDER_CREDENTIAL_CREATED / _REVOKED / _DELETEDA provider key was added or removed
CREDENTIAL_ROTATED / _INVALIDATED / CREDENTIAL_ROTATION_FAILEDA provider key was rotated, retired, or failed to rotate

WORKSPACE_SWITCH_REFUSED is the sharper half of that pair: an attempt to act outside your own team is what a reviewer looks for.

Money and licensing.

EventWritten when
BUDGET_CAP_HARDA request was refused because a budget cap was reached
BUDGET_CAP_SOFTA soft budget threshold was crossed; the request still ran
BUDGET_CAP_CREATED / _UPDATED / _DELETEDA budget cap was changed
BUDGET_COLD_START_REFUSEDA capped workspace was refused because spend history was unavailable
LICENSE_EXPIRY_WARNING / LICENSE_EXPIRED / LICENSE_DEGRADEDThe licence is approaching or past expiry

Every request.

EventWritten when
GATEWAY_REQUESTA request entered the gateway
GATEWAY_RESPONSEA request completed, carrying the policy decision, model, and token counts

How it works

  1. On the way out, the gateway writes a GATEWAY_RESPONSE audit event with a rich payload: model, provider, HTTP status, latency, tokens, masked API key, workspace ID, policy decision, and error code.
  2. Every event is both logged to stdout (for your log pipeline) and persisted to the audit store.

Query Endpoints

Browse the audit trail on the Console's Audit page (/audit), which filters by workspace, event type and date range, and expands any row to its full JSONB payload. A workspace sees its own events at /portal/audit.

Export from the same page: /audit/export returns CSV by default and JSON with ?format=json. Prefer JSON for a SIEM import or a diff — a payload is JSONB, and CSV can only carry it as one squashed cell.

Event Fields

FieldTypeDescription
eventIdstringUnique event ID (UUID)
timestampISO 8601When the event occurred
tenantIdstringWorkspace identifier (may be null)
eventTypestringEvent type — GATEWAY_RESPONSE on every data-plane request; many other types fire for policy decisions, PII / guardrail / MCP enforcement, and admin actions (e.g. POLICY_DENIED, PII_DETECTED, AGENT_LOOP_DETECTED, WORKSPACE_CREATED). See SIEM and Webhooks for the full event-type catalog.
payloadobjectEvent-specific data (model, provider, status, latency, etc.)

Storage

Audit events are persisted to PostgreSQL and indexed by workspace, event type, and timestamp. The audit store is append-only — events cannot be updated or deleted through the API.

A2A plane audit (separate store)

The A2A governance plane keeps its hop audit (A2A_HOP_INTENT / A2A_HOP_RESULT / A2A_HOP_DENIED, A2A_DELEGATION_RECORDED, A2A_PII_DETECTED, A2A_LOOP_DETECTED, A2A_CARD_DISCOVERED) on its own tamper-evident HMAC hash chain, in a store separate from the shared audit events above — so the two streams stay independently auditable. These events therefore do not appear in the shared audit viewer. Browse them at Console → Agents → A2A Audit (cross-workspace, with per-event tamper-evidence) or Portal → Agents → A2A Audit (the workspace's own hops).

They are included in the SIEM export, though — the A2A own-chain events are forwarded to Splunk / CloudWatch / Kafka alongside the shared stream, with the same source/sourcetype, so your security team sees agent-to-agent activity in the same place as everything else.

They also surface in every generated compliance report (SOC 2, HIPAA, GDPR, RBI, SEBI): each report carries an A2A governance section — hop volume, per-hop policy denials, PII actions, and loop detections for the window — plus an A2A chain-integrity check that verifies the separate A2A trail's tamper-evidence the same way it verifies the shared chain. So an auditor sees agent-to-agent activity and its forensic integrity in the same document as the LLM and MCP evidence.

Tamper-evident audit trail

The audit subsystem provides:

  • HMAC-SHA256 signing and hash-chaining. Every audit event is wrapped in a signed envelope. Each envelope's HMAC includes the previous event's hash, creating a tamper-evident chain that can be verified end-to-end.
  • Event enrichment. Events are automatically enriched with trace_id from the request context and with actor_user_id / actor_user_name / actor_roles from the authenticated principal.
  • Prompt storage opt-in. By default, prompt and message content is stripped from audit events. Workspaces opt in by setting audit.store-prompts: "true" in workspace metadata. The global default is controlled by dvara.audit.store-prompts-by-default.
  • SIEM export. Signed envelopes are fanned out to any combination of built-in exporters: a JSON log exporter (always active), Kafka (workspace-keyed partitioning, dead-letter topic, SASL support), Splunk HEC, and AWS CloudWatch Logs. Export failure never blocks audit persistence. See SIEM and Webhooks for the full exporter configuration.
  • Chain integrity verification. The whole hash chain or individual events can be verified by recomputing HMACs, either through the admin API or from a background integrity sweep.

Enterprise audit configuration:

docker run -d --name dvara-gateway \
-p 8080:8080 \
-e DVARA_LICENSE_KEY=<your-license-key> \
-e DVARA_AUDIT_HMAC_SECRET=your-hmac-secret \
-e DVARA_AUDIT_MAX_EVENTS=100000 \
-e DVARA_AUDIT_STORE_PROMPTS=false \
ghcr.io/dvarahq/dvara-gateway:1.7.0

DVARA Flightdeck

The audit log is viewable in the DVARA Flightdeck at /audit with live 3-second polling, filtering by workspace / event type / date range, click-to-expand event details, pause and resume, and CSV export.