venice-responses Skill
使用 Venice 的 Alpha POST /responses 端点——一个 OpenAI 兼容、无状态的 Responses API,带类型化输出块(reasoning、message、function_call、web_search_call)。涵盖请求结构、输入项、工具(function、web_search、x_search)、推理控制、未完成响应、流式事件、与 /chat/completions 的差异、支持的 venice_parameters 子集,以及 E2EE 行为。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Venice Responses API (Alpha)
POST /api/v1/responses is Venice's OpenAI-compatible Responses endpoint. It returns a typed output array instead of a single message.content string — useful for agents that need to separate reasoning, messages, tool calls, and web-search events. Internally the request is translated to a chat completion, so model support matches venice-chat.
Alpha. The spec labels it Alpha (and its description still says "Alpha testers only"), but access is no longer restricted: any Bearer API key or x402 wallet can call it. Schemas may still change.
Use when
- A client library expects the OpenAI Responses shape (
output[]withtype: "reasoning" | "message" | "function_call" | "web_search_call"). - You want reasoning, message, and tool-call output cleanly separated.
- You want SSE streaming with typed events.
Otherwise use venice-chat — it has structured output, audio/video/file inputs, E2EE, sampling controls, and every venice_parameters field.
Limitations vs /chat/completions
| Limitation | Detail |
|---|---|
| Stateless | Nothing is stored. Send the full history each call. previous_response_id, store, background are ignored. |
| No E2EE | E2EE-capable models return 400 unless venice_parameters.enable_e2ee: false (TEE-only mode). For encrypted inference use /chat/completions. |
| Text + image input only | input_text / input_image (and text / image_url parts). No audio, video, or file parts. |
| No structured output | text.format / response_format are dropped. Use /chat/completions. |
| Subset of venice_parameters | character_slug, enable_e2ee, enable_web_search, enable_web_scraping, enable_web_citations, include_venice_system_prompt, include_search_results_in_stream. Other keys (strip_thinking_response, disable_thinking, enable_x_search, return_search_results_as_documents) are silently dropped. |
| No model feature suffixes | model: "zai-org-glm-5-1:enable_web_search=on" resolves the model but ignores the suffix. |
| Few generation controls | Only temperature, top_p, max_output_tokens. |
Authentication
Same as the rest of the API — Authorization: Bearer or SIGN-IN-WITH-X: for x402 wallets. See venice-auth.
Minimal request
curl https://api.venice.ai/api/v1/responses \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zai-org-glm-5-1",
"input": "Explain why the sky is blue in one paragraph."
}'
OpenAI SDK (Python):
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["VENICE_API_KEY"], base_url="https://api.venice.ai/api/v1")
resp = client.responses.create(model="zai-org-glm-5-1", input="Explain why the sky is blue.")
print(resp.output_text)
Request fields
| Field | Notes |
|---|---|
| model | Required. Model ID, trait, or compatibility mapping. |
| input | Required. A string, or an array of input items (below). |
| max_output_tokens | Positive integer. Mapped to max_tokens; above the model's model_spec.maxCompletionTokens → 400 on models with an enforced cap. |
| temperature (0–2), top_p (0–1) | Sampling. |
| reasoning.effort | none \| minimal \| low \| medium \| high \| xhigh \| max (per-model support: model_spec.capabilities.reasoningEffortOptions). reasoning may be null. |
| reasoning.enabled | false disables reasoning on supported models and suppresses reasoning blocks. Ignored when an effort is set. |
| reasoning.summary | auto \| concise \| detailed. Accepted but not forwarded. |
| tools | See Tools. |
| tool_choice | "auto" \| "none" \| "required" \| {"type":"function","function":{"name":"..."}}. Dropped when no function tools remain. |
| web_search | Boolean. true forces web search on (same as a {"type":"web_search"} tool). |
| include | Only "reasoning.encrypted_content" has an effect (adds encrypted_content to reasoning blocks when the provider returns it). |
| stream | Boolean. SSE with typed events. |
| anon_user_id | Optional end-user id: 1–128 printable ASCII characters, no ||. Distinct from user. |
| fallbacks | Up to 10 {model} entries. Anthropic beta refusal fallback for Claude Fable 5; forwarded only on direct Anthropic routes. |
| venice_parameters | Subset listed above. Example: {"character_slug":"alan-watts","enable_web_search":"auto"}. |
The body is permissive: other fields (instructions, metadata, parallel_tool_calls, n, stop, seed, prompt_cache_key, store, previous_response_id, background, text, user) are accepted without error but never reach inference (user still splits the error budget per value). Put system instructions in the input array instead of instructions.
Input items
| Item | Shape | Handling |
|---|---|---|
| Message | {role, content} or {type:"message", role, content} | role: user / assistant / system / developer (developer becomes system). content is a string or an array of parts. |
| Function call | {type:"function_call", call_id, name, arguments} | Replayed as an assistant tool call. |
| Function output | {type:"function_call_output", call_id, output} | output may be a string, array, object, number, boolean, or null. input_image parts inside an array output are forwarded to the model as images. |
| Reasoning | {type:"reasoning", ...} | Accepted but discarded — reasoning is not carried between turns. |
| Item reference | {type:"item_reference", id} | Accepted but discarded (nothing is stored to reference). |
Content parts: input_text, output_text (to replay assistant output), and input_image. input_image.image_url may be a URL string (OpenAI Responses style) or {url, detail}; detail (auto / low / high) may also sit on the part. Messages without type additionally accept Chat-style text and image_url parts. Image URLs get the same validation as on /chat/completions (public, no redirects, ≥ 64 px); failures → 400. Use a vision model: message images are not capability-checked on this endpoint (images inside a function_call_output on a non-vision model do return 400).
Tools
| Tool | Effect |
|---|---|
| {"type":"function","function":{name, description, parameters, strict}} | Function calling. The flat OpenAI form {"type":"function","name":...,"parameters":...} is also accepted. Use a model with supportsFunctionCalling (not pre-checked on this endpoint, unlike chat). |
| {"type":"web_search"} | Forces Venice web search on (not auto). search_context_size / user_location are accepted but ignored. |
| {"type":"x_search", ...} | xAI native web + X search on models with supportsXSearch (Grok); ignored on other models. Optional filters: allowed_x_handles / excluded_x_handles (≤ 10 each), from_date, to_date, enable_image_understanding, enable_video_understanding. |
| code_interpreter, file_search, computer_use_preview, others | Accepted and dropped. Unknown tool types that carry a name are treated as function tools. |
Response shape
{
"id": "resp_chatcmpl-abc123",
"object": "response",
"created_at": 1735689600,
"model": "zai-org-glm-5-1",
"status": "completed",
"output": [
{"type": "reasoning", "id": "rs_1", "summary": ["I considered Rayleigh scattering..."]},
{"type": "web_search_call", "id": "ws_1", "status": "completed"},
{"type": "function_call", "id": "fc_1", "call_id": "call_abc", "name": "get_weather",
"arguments": "{\"city\":\"Paris\"}", "status": "completed"},
{"type": "message", "id": "msg_1", "status": "completed", "role": "assistant",
"content": [{"type": "output_text", "text": "The sky is blue because... ^1^",
"annotations": [{"type": "url_citation", "url": "https://example.com/rayleigh",
"title": "Rayleigh scattering", "start_index": 27, "end_index": 30}]}]}
],
"usage": {
"input_tokens": 20,
"input_tokens_details": {"cached_tokens": 8},
"output_tokens": 80,
"output_tokens_details": {"reasoning_tokens": 40},
"total_tokens": 100
}
}
- Non-streamed
outputorder:reasoning→web_search_call→function_call(s) →message. Themessageblock is omitted when the model returned only tool calls with no text. statusiscompletedorincomplete. Errors before or during inference come back as HTTP errors (streaming usesresponse.failed).- Incomplete responses keep their output and usage. When generation stops on
max_output_tokensor a content filter,status: "incomplete",incomplete_details: {"reason": "max_output_tokens" | "content_filter"}, and message / function_call blocks carrystatus: "incomplete". input_tokens_detailsappears only when cached tokens are non-zero;output_tokens_detailsonly when the provider reports reasoning tokens. There is nocostfield (unlike/chat/completions).
Output block types
| type | Purpose |
|---|---|
| reasoning | Reasoning from thinking models. summary[] holds text; encrypted_content appears only if you sent include: ["reasoning.encrypted_content"] and the provider returned encrypted reasoning. Sending it back in input has no effect. |
| message | Main text. content[].type === "output_text" with annotations[]. |
| function_call | Tool call: name, JSON-string arguments, call_id. Answer with a function_call_output item with the same call_id. |
| web_search_call | Marker that Venice web search ran. |
url_citation annotations are built only when the text contains single-index ^n^ markers — set venice_parameters.enable_web_citations: true to get them. Each annotation spans the marker itself; multi-index markers such as ^1,3^ are not annotated.
Streaming
With stream: true, events are event: + data: {...} pairs; payloads carry type (equal to the event name) and an increasing sequence_number — except response.web_search.done, whose payload has type: "web_search_call", id, status, results and no sequence_number. Typical flow:
event: response.created # status: in_progress
event: response.web_search.done # Venice-specific; only when search ran, carries results[{index,url,title,snippet}]
event: response.output_item.added # item.type = reasoning
event: response.reasoning.delta
event: response.output_item.added # item.type = message
event: response.content_part.added
event: response.output_text.delta # repeated
event: response.output_item.added # item.type = function_call
event: response.function_call_arguments.delta
event: response.output_item.done # reasoning, then message (after content_part.done), then each function_call
event: response.completed # or response.incomplete, with the full response
data: [DONE]
- On an upstream failure you get
response.failed(response.status: "failed",response.error: {code, message}) followed bydata: [DONE]. - The final
response.completed/response.incompletepayload differs slightly from a non-streamed response:annotationsare always empty and function calls come after the message. include_search_results_in_streamhas no effect here; search results always arrive inresponse.web_search.done.
Errors
| Status | When |
|---|---|
| 400 | Invalid body, E2EE-capable model without enable_e2ee: false, invalid image, unsupported reasoning.effort for the model, context too long, max_output_tokens over the cap |
| 401 | Invalid API key or SIWX sign-in; also a model that requires a paid subscription |
| 402 | No credentials at all (x402 discovery body — not 401), insufficient balance, or API-key spend limit. x402: PAYMENT_REQUIRED body with topUpInstructions + siwxChallenge and a PAYMENT-REQUIRED header (see venice-x402) |
| 403 | Model blocked by the key's modelPrivacy, region, or provider restriction |
| 404 | Unknown model |
| 422 | Content-policy violation on an input image |
| 429 | Rate limited |
| 500 / 503 | Inference failed (upstream overloads and timeouts also surface as 500 here, not 429 / 504; retry with backoff) / model offline |
The spec lists an X-Balance-Remaining header on x402 200 responses, but the server does not currently set it — poll GET /x402/balance/{walletAddress} instead. See venice-errors.
Migration notes (from /chat/completions)
messages→input(the same role/content objects work; system prompts go in asrole: "system"or"developer"items).max_tokens→max_output_tokens;reasoning_effort→reasoning.effort.- Tool results →
function_call_outputitems keyed bycall_id. venice_parameters.character_slug,enable_web_search,enable_web_citations,enable_web_scraping,include_venice_system_prompt→ pass insidevenice_parameters(not as model suffixes).enable_x_search→ add an{"type":"x_search"}tool instead.strip_thinking_response/disable_thinking→ usereasoning.enabled: false.- Structured output, audio / video / file inputs,
seed/stop/n, logprobs, prompt-cache routing, and full E2EE → stay on/chat/completions.
Gotchas
- Unknown
character_slugis not rejected here (chat returns404). The request runs without the character, without yoursystemmessages, and without the Venice system prompt or web search. Validate slugs first withGET /characters/{slug}(venice-characters; Bearer key only — wallet callers can't, and should use/chat/completions, which returns404for unknown slugs). - Reasoning items in
inputare discarded; there is no cross-turn reasoning carry-over on this endpoint. tool_choiceobjects must be{"type":"function","function":{"name":...}}; the flat{"type":"function","name":...}form fails validation.- Stateless:
previous_response_idis silently ignored, so omitting history silently loses context.