Skip to main content
Version: Latest (1.9.x dev)

Function Calling (Tools)

DVARA is an AI governance platform for model and tool traffic. On POST /v1/chat/completions, your application can send OpenAI-compatible function definitions, receive the model's tool calls, execute them, and return the results while DVARA applies the configured controls to every turn.

DVARA does not execute the function or run an agent loop. Your application owns those steps. The DVARA LLM Gateway translates provider formats, governs the model call, and returns one OpenAI-compatible choice.

Send a function definition​

Replace <your-dvara-api-key> and the hostname with values for your deployment.

curl -s https://dvara.internal.example.com/v1/chat/completions \
-H 'Authorization: Bearer <your-dvara-api-key>' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
],
"tool_choice": "auto"
}'

Only entries with a function definition are relayed. Hosted and built-in tool types are not supported on this endpoint.

Read the model's tool call​

When the model chooses the function, the assistant message carries the call and the choice ends with finish_reason: "tool_calls":

{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_1",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\":\"Paris\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
]
}

function.arguments is a JSON string produced by the model. Parse it only after you have received the complete call, then validate it before executing the function.

Return the tool result​

Execute the function in your application. On the next request, send the full history: the original messages, the assistant message with tool_calls, and a tool message whose tool_call_id matches the call.

curl -s https://dvara.internal.example.com/v1/chat/completions \
-H 'Authorization: Bearer <your-dvara-api-key>' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "What is the weather in Paris?"},
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_1",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\":\"Paris\"}"
}
}
]
},
{
"role": "tool",
"tool_call_id": "call_1",
"content": "Sunny, 24 C"
}
]
}'

The model can now answer using the result:

{
"id": "chatcmpl-def456",
"object": "chat.completion",
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "It is sunny and 24 C in Paris."
},
"finish_reason": "stop"
}
]
}

An assistant message can carry tool_calls with no text content. A message can also include the optional OpenAI-compatible name field.

Route only to a capable provider​

When a request includes tools, DVARA removes providers that do not declare native tool-call support. When it also sets stream: true, DVARA removes providers that cannot preserve tool calls on a stream. If the route has no capable provider, DVARA refuses the request instead of silently dropping the function definition or call.

GET /v1/models exposes supports_tool_calls, but it does not separately expose streamed-tool support in 1.8. Use the provider capabilities matrix for both columns. DVARA translates the OpenAI-compatible request to each provider's native shape.

Stream native tool calls safely​

Set stream: true on the first request. The opening fragment identifies the call and function; later fragments carry slices of its JSON argument string:

data:{"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"gpt-4o","choices":[{"index":0,"delta":{"role":"assistant","tool_calls":[{"index":0,"id":"call_1","type":"function","function":{"name":"get_weather","arguments":""}}]},"finish_reason":null}]}

data:{"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"gpt-4o","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\":"}}]},"finish_reason":null}]}

data:{"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"gpt-4o","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"Paris\"}"}}]},"finish_reason":null}]}

data:{"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"gpt-4o","choices":[{"index":0,"delta":{},"finish_reason":"tool_calls"}]}

data:[DONE]

Keep one argument buffer per index, append fragments in arrival order, and parse each buffer only after the terminal tool_calls finish reason. Calls can be interleaved, so never concatenate fragments from different indexes.

DVARA also assembles the fragments by index for response governance. It parses each completed JSON value and scans each string and number separately. Under a withholding action, DVARA holds the response until it can govern every complete call, then releases one complete, valid JSON argument string per call.

Deferred delivery refuses unchecked arguments. Invalid JSON, duplicate keys, more than 512 argument values, or an edit that cannot be rendered safely produces STREAM_TOOL_ARGUMENTS_UNENFORCEABLE. More than 256 calls or the configured held-response character limit produces STREAM_TOO_LARGE_TO_SCAN. Because the HTTP stream is already open, the client receives a terminal chunk with finish_reason: "content_filter"; the specific reason is recorded in audit evidence.

With the default PII LOG and guardrail LOG actions, delivery is Immediate: DVARA relays each fragment as it arrives and scans the assembled arguments at the end. If the JSON cannot be scanned safely, it records scan_incomplete in STREAMING_ENFORCEMENT_SUMMARY; it cannot retract fragments already sent. Configure a withholding action such as PII BLOCK or REDACT when unchecked arguments must never reach the client. See SSE streaming for the complete delivery contract and defaults.

Apply governance to both turns​

Model-produced tool_calls[].function.arguments are response content. DVARA applies the configured response PII and guardrail controls before Deferred delivery, or records findings after Immediate delivery. Tool-call arguments replayed inside an assistant message and tool results sent with role: "tool" are request content on the next turn and pass through the configured request controls.

PII defaults to LOG. BLOCK refuses the content, REDACT replaces detected values irreversibly, and Enterprise TOKENIZE creates recoverable replacements on the request side. Guardrails support LOG, FLAG, and BLOCK. The behavior you see therefore depends on the workspace actions you configure.

Verify audit and usage​

Every governed stream writes STREAMING_ENFORCEMENT_SUMMARY when audit storage is configured. Findings add their corresponding PII or guardrail event, and a refused argument payload adds STREAM_TOOL_ARGUMENTS_UNENFORCEABLE. The normal completed-request evidence also includes GATEWAY_RESPONSE.

Usage includes generated text plus function names and argument fragments. DVARA uses the provider's terminal usage when it is available; otherwise it records an estimate and marks the usage row estimated. The public SSE stream does not add a DVARA-specific usage object.

Know the unsupported paths​

Function calling is supported on /v1/chat/completions only. The /v1/responses endpoint refuses tools and tool_choice with UNSUPPORTED_CAPABILITY instead of dropping them.

Hosted and built-in tools such as web search, file search, code interpreter, computer use, and remote MCP are not relayed. To govern MCP tool traffic, use the DVARA MCP Gateway.