venice-errors Skill
正确处理 Venice API 错误。涵盖错误体结构(StandardError、带验证详情/问题的 DetailedError、OpenAI 风格的 context_length_exceeded、上游供应商拒绝及含 request_id 的 TypeSafe/FastAPI 验证错误、ContentViolationError、ProviderContentPolicyError、PayloadTooLargeError、两种 x402 402 错误体)、所有有意义的状态码(400-504,含已下线端点的 410)、三类限流头、按 key 计的错误预算(FAILED_REQUESTS / UNSUPPORTED_FEATURE_REQUESTS)、过载模型的 Retry-After,以及 SSE 流错误
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Venice errors & retries
Most Venice errors are { "error": "" }, but several paths add structure. Knowing which shape you got tells you how to react.
Error body shapes
1. StandardError - simple message
The default for 4xx/5xx.
{ "error": "Authentication failed" }
2. DetailedError - schema validation failure (400)
When a request fails Venice's own schema, details is a nested tree (_errors recursively keyed by field) and issues is the flat issue list. Some image, video and upstream failures instead return details as a plain string, so check its type before walking it.
{
"error": "Invalid request parameters",
"details": {
"_errors": [],
"type": { "_errors": ["Invalid enum value. Expected 'asr' | 'decision' | … , received 'bogus'", "Invalid enum value. Expected 'all' | 'code', received 'bogus'"] }
},
"issues": [
{ "code": "invalid_union", "unionErrors": [ … ], "path": ["type"], "message": "Invalid input" }
]
}
(That is the live response to GET /models?type=bogus.)
Many 400s are plain StandardError - handle both. Render details / issues; don't retry.
3. OpenAI-style context overflow (400, chat)
{
"error": {
"message": "Your request exceeds the model's maximum context. Please reduce your prompt or completion length.",
"type": "invalid_request_error",
"param": "messages",
"code": "context_length_exceeded"
}
}
Note error is an object here, so OpenAI SDKs can detect it. message is the provider's wording when Venice can extract it, otherwise the text above - match on code, not message. Trim the prompt or lower max_completion_tokens.
4. Upstream provider rejection (400)
When the request passed Venice's schema but the model provider rejected it, Venice returns the provider's message, plus a request_id on routes that track one (chat, for example; /decisions omits it):
{ "error": "<provider's validation message>", "request_id": "…" }
FastAPI-style validation errors (e.g. from the TypeSafe provider behind POST /decisions, which answers 422) are surfaced as 400 with up to five field.path: message items joined by ; . If Venice can't extract a message you get "Invalid request parameters. For assistance, please reach out to support@venice.ai", followed by " and reference request ID: " when there is one. These count against the lenient unsupported-feature budget on chat / responses (see Error budget). Fix the input; quote request_id to support.
5. ContentViolationError - 422 content policy
{
"error": "Your prompt violates the content policy of Venice.ai or the model provider",
"suggested_prompt": "A cinematic instrumental track inspired by stormy weather and dramatic tension."
}
Returned by chat, responses, image edit and multi-edit, and audio generation. /image/generate does not use it (a blocked image comes back as a 200 with x-venice-is-content-violation: true), and video content-policy rejections arrive on /video/retrieve rather than on queue. suggested_prompt is only emitted by /audio/queue and /audio/retrieve; when present, retry once with it if the user consents. Other 422s include "Your input was blocked by content moderation." and media-validation failures (image too large, bad aspect ratio, audio/video duration out of range, ASR unable to process the audio).
6. ProviderContentPolicyError - 422 on /video/retrieve
{
"error": {
"message": "The selected model provider rejected this request due to its content policies. Credits have been refunded. Try using <model> instead.",
"type": "provider_content_policy",
"credits_refunded": true,
"recommended_model": "<model id>"
}
}
Check credits_refunded; optionally re-queue on recommended_model.
7. PayloadTooLargeError - 413
{ "code": "PAYLOAD_TOO_LARGE", "error": "File exceeds the maximum allowed size of 25 MB." }
8. x402 402 bodies
A 402 comes in two forms. Both also set the PAYMENT-REQUIRED header. The no-credentials form is returned on every route that needs credentials, including Bearer-only routes (/api_keys, /billing/, /characters*), where its payment options don't apply.
No credentials at all - x402 v2 discovery (not 401):
{
"x402Version": 2,
"error": "Authentication required",
"resource": { "url": "https://api.venice.ai/api/v1/chat/completions", "description": "Venice API", "mimeType": "application/json" },
"accepts": [ { "scheme": "exact", "network": "eip155:8453", "…": "…" }, { "scheme": "exact", "network": "solana", "…": "…" } ],
"extensions": { "sign-in-with-x": { "info": { "domain": "api.venice.ai", "statement": "Sign in to Venice AI", "…": "…" }, "supportedChains": [ … ] } },
"authOptions": {
"apiKey": { "header": "Authorization: Bearer YOUR_API_KEY", "getKey": "https://venice.ai/settings/api", "docs": "…" },
"x402Wallet": { "header": "SIGN-IN-WITH-X", "legacyHeader": "X-Sign-In-With-X", "topUp": "POST /api/v1/x402/top-up", "docs": "…" }
}
}
Signed-in wallet below the $0.10 minimum balance - discriminate on code: "PAYMENT_REQUIRED":
{
"error": "Payment required",
"code": "PAYMENT_REQUIRED",
"reason": "insufficient_balance",
"currentBalanceUsd": 0.01,
"minimumBalanceUsd": 0.1,
"description": "Venice API",
"suggestedTopUpUsd": 10,
"minimumTopUpUsd": 5,
"supportedTokens": ["USDC"],
"supportedChains": ["base", "solana"],
"topUpInstructions": {
"step1": "POST /api/v1/x402/top-up with no payment header to get payment requirements",
"step2": "Choose a payment option from accepts and sign a USDC transfer authorization using the x402 SDK (createPaymentHeader)",
"step3": "POST /api/v1/x402/top-up with the signed X-402-Payment header",
"receiverWallet": "0x…",
"tokenAddress": "0x…",
"tokenDecimals": 6,
"network": "eip155:8453",
"minimumAmountUsd": 5
},
"siwxChallenge": { "info": { … }, "supportedChains": [ … ] }
}
topUpInstructions describes the Base rail only; read accepts[] from POST /x402/top-up to pay on Solana (and send PAYMENT-SIGNATURE, the canonical header - X-402-Payment / X-PAYMENT still work). The PAYMENT-REQUIRED header is the base64 x402 paymentRequired object (x402Version, error, resource, accepts[], extensions), not the body. See venice-x402.
API-key 402s are plain: { "error": "Insufficient USD or Diem balance to complete request. Visit https://venice.ai/settings/api to add credits." }, or the per-key spend-limit variants ("API key DIEM spend limit exceeded…" / "API key USD spend limit exceeded…"). A wallet can get the same plain "Insufficient USD or Diem balance…" body when its credit clears the $0.10 floor but not the quoted price of this request (e.g. /video/queue, /audio/queue) - top up and retry.
9. x402 sign-in failures (401)
On inference routes a bad SIGN-IN-WITH-X returns { "error": "", "code": "X402_SIGN_IN_…" } (e.g. X402_SIGN_IN_EXPIRED, X402_SIGN_IN_NONCE_REUSED). /x402/balance and /x402/transactions return a generic { "error": "Invalid Sign-in-with-x signature" }. See venice-auth for every code.
Status code map
| Status | Typical body | Meaning | What to do |
|---|---|---|---|
| 400 | DetailedError, StandardError, context-overflow object, or upstream { error, request_id } | Malformed input, missing or non-string model (plain "model is required" / "model must be a string"), unsupported option for this model, provider rejection, invalid JSON ("Invalid JSON request"), or a POST whose Content-Type is neither JSON nor multipart ("'Content-Type' must be 'application/json'"). Also PAYMENT_HEADER_NOT_ACCEPTED if you send an x402 payment header to an inference route instead of /x402/top-up. | Fix and re-send. Don't retry. |
| 401 | StandardError or { error, code } | Unknown, expired or revoked API key ("Authentication failed"), a non-ADMIN key on an admin-only route ("Admin API key required"), bad SIWX, or "This model is only available to Pro users" (API-key accounts without a paid plan on a Pro-only model). | Fix credentials / plan. Don't retry. |
| 402 | See shape 8 | No credentials (discovery), wallet balance too low, or API-key account / key spend limit exhausted. | x402: top up then retry. API key: add credits or raise the key's limit. |
| 403 | StandardError | Entitled-but-blocked: model blocked in your country (regionRestrictions), key's modelPrivacy forbids the model, API access disabled, SIWX wallet ≠ path wallet. | Don't retry. |
| 404 | StandardError | Unknown model ("Specified model not found: …", sometimes with a suggestion or a "has been deprecated. Please use …" hint), unknown character, expired media. | Fix the ID. |
| 409 | { error: { code: "needs_consent", message }, consent_flow, face_media_roles, consent, docs_url } (video) or { error, message, details } (x402) | /video/queue needs consent (only on unlisted model ids; listed models never ask), or an x402 top-up that is already processed or still settling (error holds the code, e.g. "PAYMENT_IN_PROGRESS" — retry shortly). | See venice-video / venice-x402. |
| 410 | StandardError | Retired endpoint (GET /billing/usage, POST /video/transcriptions). Message names the replacement. | Migrate. Never retry. |
| 413 | PayloadTooLargeError | JSON body over 35 MB ("Request body exceeds the maximum allowed size.") or a multipart file over 25 MB. | Shrink the upload. |
| 415 | StandardError | Rare: /image/multi-edit with an empty body, or a body sent with a Content-Encoding / charset the server can't decode ("Request encoding is not supported"). A wrong Content-Type on any route is answered with 400. | Fix headers. |
| 422 | ContentViolationError, ProviderContentPolicyError, or StandardError | Content policy, or media that can't be processed (dimensions, duration, unreadable audio). | Change the prompt / media. One retry with suggested_prompt if offered. |
| 429 | StandardError or { error, code } | Request/token rate limit, error budget exhausted, model overloaded, x402 concurrency (X402_CONCURRENCY_LIMIT, 5 in-flight per wallet), or a per-route limiter (crypto RPC, /x402/, /tee/, retired endpoints). | See rate limits below. |
| 500 | StandardError | Unexpected failure. | Backoff and retry. |
| 502 | StandardError | Upstream failure (TTS, ASR, TEE, video fetch). | Backoff and retry. |
| 503 | StandardError | Model offline or at capacity. | Backoff; consider a fallback model. |
| 504 | StandardError | Request took too long. Mostly non-streaming chat. | Use stream: true or a smaller request. |
Rate limits and their headers
Three independent header families - they share a prefix but not units, so read them by exact name:
| Headers | Emitted by | Reset unit |
|---|---|---|
| x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests (+ -tokens variants) | Model rate limits on inference routes (per account, per model, requests per minute/day and tokens per minute) | Unix milliseconds |
| x-ratelimit-remaining, x-ratelimit-resets | The error budget, on most responses from routes that need credentials (not on the no-credentials 402) | Unix milliseconds |
| X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset | /crypto/rpc/{network}, only on the 429 from its per-minute cap | Unix seconds |
Model-limit 429s say "Rate limit exceeded". Pre-fetch your caps with GET /api_keys/rate_limits (Bearer keys only). Past hits are in GET /api_keys/rate_limits/log, which needs an ADMIN key — an INFERENCE key gets 401 "Admin API key required" (venice-api-keys).
Per-route limiters return 429 without these headers: /x402/top-up (10/min per IP), /x402/balance (30/min per wallet) and /x402/transactions (20/min per wallet) say "Rate limit exceeded. Please try again later."; /tee/* (10/min per IP) says "Rate limit exceeded."; the retired endpoints allow 60/min per IP.
Overloaded upstream: 429 with Retry-After (seconds, default 30) and a message like "The model is currently overloaded. Please try again later." Honor Retry-After.
Error budget
Failed requests are rate-limited harder than successful ones:
| Budget | Threshold | Counts | Logged as |
|---|---|---|---|
| Failed requests | 50 per 30 s | Any 4xx except 429 (5xx never count) | FAILED_REQUESTS |
| Unsupported feature requests | 200 per 30 s, /chat/completions and /responses only | Requests asking a model for a capability it lacks, or provider-side rejections of otherwise valid requests | UNSUPPORTED_FEATURE_REQUESTS |
Buckets are per API key (or per IP for x402), per model, per OpenAI user string; multipart uploads don't expose model / user in time, so they share the key's (or IP's) bucket. Once a bucket is exhausted every request in it gets 429 until reset, with:
Too many failed attempts (> 50) resulting in a non-success status code. Please wait 30 seconds and try again. See https://docs.venice.ai/api-reference/rate-limiting for more information.
The unsupported-feature bucket uses the same wording with > 200. So a client that blindly retries 400/401/402 locks itself out. Stop on non-retryable errors. (The no-credentials 402, the Content-Type and invalid-JSON 400s, and the 35 MB JSON-body 413 are rejected before the budget is checked and don't count.)
Retry strategy
Never retry
400, 401, 403, 404, 410, 413, 415 - fix the request, credentials, or endpoint.
Retry with modification
402withcode: "PAYMENT_REQUIRED"- top up via/x402/top-upwithin the user's spend cap (seevenice-x402), then retry.402withx402Version(no credentials) - addAuthorization: Bearer …, orSIGN-IN-WITH-Xon routes that accept wallets. Bearer-only routes (/api_keys,/billing/,/characters*) reject SIWX with401; send a Bearer key there.- Any retry with wallet auth needs a newly signed
SIGN-IN-WITH-Xheader - each nonce is accepted once, so re-sending the original header fails withX402_SIGN_IN_NONCE_REUSED. 402on an API key - surface to the user.422withsuggested_prompt- one retry with the safer prompt.
Retry with backoff
429- waitRetry-Afterif present, otherwise until the relevant reset header; add jitter.500/502/503/504- exponential backoff (0.5 s, 1 s, 2 s, 4 s, 8 s), capped at ~30 s, 3-5 retries max.- Async queues (
/video/queue,/audio/queue,/audio/voice-changer/queue) can bill before the job finishes (API-key calls are pre-charged at queue time). If a queue response is lost, pollretrieveinstead of re-queueing.Idempotency-Keyis supported only on/crypto/rpc/{network}.
Reference retry loop
const sleep = (ms: number) => new Promise(r => setTimeout(r, ms))
function waitMs(res: Response, fallback: number): number {
const retryAfter = Number(res.headers.get('retry-after'))
if (retryAfter > 0) return retryAfter * 1000
const msReset = Number(res.headers.get('x-ratelimit-reset-requests') ?? res.headers.get('x-ratelimit-resets'))
if (msReset > 0) return Math.max(msReset - Date.now(), fallback)
const secReset = Number(res.headers.get('x-ratelimit-reset')) // crypto RPC
if (secReset > 0) return Math.max(secReset * 1000 - Date.now(), fallback)
return fallback
}
// fn must build a fresh SIGN-IN-WITH-X header on every call; nonces are single-use.
async function callVenice<T>(fn: () => Promise<Response>): Promise<T> {
const maxRetries = 5
let delay = 500
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const res = await fn()
if (res.ok) return res.json() as Promise<T>
const body = await res.clone().json().catch(() => ({}))
const message = typeof body.error === 'string' ? body.error : body.error?.message ?? 'Venice error'
const fail = () => Object.assign(new Error(message), { status: res.status, body })
if ([400, 401, 403, 404, 410, 413, 415, 422].includes(res.status)) throw fail()
if (res.status === 402) {
// topUpAllowed must enforce the user's cap and check GET /x402/balance first.
if (body.code === 'PAYMENT_REQUIRED' && attempt === 0 && (await topUpAllowed())) {
await topUpX402(Math.min(body.suggestedTopUpUsd, USER_TOP_UP_CAP_USD))
continue
}
throw fail()
}
if ((res.status === 429 || res.status >= 500) && attempt < maxRetries) {
await sleep((res.status === 429 ? waitMs(res, delay) : delay) + Math.random() * 250)
delay = Math.min(delay * 2, 30_000)
continue
}
throw fail()
}
throw new Error('Exceeded max retries')
}
Streaming errors
If a /chat/completions stream fails after headers are sent, the HTTP status stays 200 and the error arrives in-band, followed by the terminator:
data: {"error":{"message":"…","type":"server_error","code":"model_overloaded","param":null,"retry_after":30}}
data: [DONE]
Overload errors use type: "server_error", code: "model_overloaded" and may carry retry_after (seconds); other upstream failures use type: "api_error" with code: "upstream_error" or a specific code (e.g. e2ee_attestation_stale). Treat the event as terminal.
Request-ID correlation
Upstream-rejection bodies on chat carry request_id, and /crypto/rpc/{network} sets an X-Request-ID header on proxied responses. Include either (plus x-venice-version from the response headers) in support tickets. Other routes don't guarantee a request ID, so keep your own client-side correlation ID too.
Common gotchas
- A
402from/x402/top-upwith no payment header is the expected discovery response. - A
402(not401) on an inference route usually means you sent no auth header at all. x-ratelimit-remainingwithout a suffix is the error budget, not your request quota.- A
429can come from several different limiters (model limits, error budget, overload, x402 concurrency, per-route caps) - read the message and headers before deciding how long to wait. DetailedError.detailsis a nested_errorstree, not a flat map, on schema-validation400s; some image / video / upstream failures send it as a plain string. Check its type before walking it.- In SSE streams,
data: [DONE]is the end of the stream (also after an error chunk), not a keepalive.