Skip to main content
Version: 1.8.0

Create a model response

POST 

/v1/responses

OpenAI-compatible Responses API. DVARA translates the request into its model-call contract and applies the same configured policy, PII, guardrail, budget, routing, usage, cost, and audit controls as Chat Completions. Non-streaming requests can use the response cache; streams bypass it. Set stream: true for typed SSE events. Observational response controls preserve incremental delivery, while a control that can withhold output holds the complete response before release or refusal. The public stream does not include usage; DVARA records terminal provider usage when available and estimates it otherwise.

Only the text + image-input + structured-output core of the Responses shape is honored in 1.8.0. Advanced OpenAI-only features are rejected cleanly with UNSUPPORTED_CAPABILITY (HTTP 400) — never silently dropped:

Rejected fieldWhy
store: true, previous_response_idThe gateway is stateless — it stores no server-side conversation state. Send the full input each turn with store: false.
background: trueAsync/background mode needs server-side state.
reasoningReasoning items are not surfaced in this release.
prompt (reusable prompt object)Use DVARA prompt templates instead.
tools, tool_choiceFunction calling on /v1/responses arrives via a follow-up; built-in hosted tools (web_search, file_search, code_interpreter, computer_use, mcp) are not supported.
input parts of type input_file / input_audioText + image input only in 1.2.0.

Both /v1/responses and /v1/chat/completions are stateless — send the full input on each turn. Function calling is supported on Chat Completions. Requests using tools or tool_choice on Responses are rejected as described above.

Request​

Responses​

Response object, or an SSE stream of typed events when stream=true