venice-audio-speech Skill
通过 POST /audio/speech 将文本转为语音,通过 POST /audio/voices 克隆声音。覆盖 TTS 模型(Kokoro、Qwen 3、xAI、Inworld、Chatterbox、Orpheus、ElevenLabs Turbo、MiniMax、Gemini Flash、Gradium)、各模型可用声音、克隆声音句柄与原始 ElevenLabs Voice ID、各模型输出格式(model_spec.supported_formats / default_format)、流式输出、提示词/风格控制、temperature/top_p、语速限制和语言提示。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Venice TTS (/audio/speech)
POST /api/v1/audio/speech converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's audio.speech.create() works as a drop-in.
| Method | Path | Auth | Notes |
|---|---|---|---|
| POST | /api/v1/audio/speech | Bearer key or x402 (SIWX) | JSON body, returns raw audio. Billed per input character. |
| POST | /api/v1/audio/voices | Bearer key or x402 (SIWX) | multipart/form-data. Clone a voice → vv_… handle. There is no GET /audio/voices. |
Use when
- You want narration, voice replies, or UI audio from text.
- You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash, Gradium).
- You want streaming audio returned as it is generated.
- You need style/emotion control on supported models, or synthesis in a cloned voice.
For music, sound effects and the character-priced ElevenLabs TTS models — v3, v4, v4 Turbo, Multilingual v2 — (async), see venice-audio-music. For transcription (audio → text), see venice-audio-transcription.
Minimal request
curl https://api.venice.ai/api/v1/audio/speech \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-xai-v1",
"voice": "eve",
"input": "Hello, welcome to Venice Voice.",
"response_format": "mp3",
"speed": 1.0,
"streaming": false
}' --output hello.mp3
Response is the raw audio. Content-Type is audio/mpeg, audio/opus, audio/aac, audio/flac, audio/wav or audio/pcm depending on the format actually produced.
Request schema
The body is strict — unknown fields return 400.
| Field | Type | Default | Notes |
|---|---|---|---|
| input | string | — | Required. 1–4096 characters, must contain non-whitespace. Text that sanitizes to nothing speakable (e.g. only markup/emoji) → 400 "Input must contain speakable text". |
| model | string | — | Send it. The OpenAPI schema lists a tts-kokoro default, but that default is never applied: omitting model returns 400 "model is required". Unknown id → 404. |
| voice | string, ≤ 512 | the model's default voice | Voices are model-specific; a voice from another model → 400. Also accepts a cloned-voice handle (vv_…) from POST /audio/voices (same model that created it), and — on models with supports_custom_voice_id: true (currently tts-elevenlabs-turbo-v2-5) — a raw provider Voice ID. |
| response_format | mp3 / opus / aac / flac / wav / pcm | the model's default_format | Support is per model — read model_spec.supported_formats / default_format. Requesting a format the model doesn't support → 400. |
| speed | number | 1.0 | Schema range 0.25–4.0. Passed to Kokoro unchanged; clamped by xAI (0.7–1.5), ElevenLabs Turbo (0.7–1.2) and MiniMax (0.5–2); ignored by the other models. |
| streaming | bool | false | true → chunked audio stream as it's generated. false → buffered file with Content-Length. |
| language | string, 2–32 chars | — | Optional hint; form is model-specific (see below). Gemini Flash drops values outside its locale list and ElevenLabs Turbo drops values longer than 5 chars; Qwen 3, MiniMax and xAI forward the value as given, and a value the model rejects returns 400 "Invalid request parameters: language". Other models ignore it. |
| prompt | string, ≤ 500 | — | Style/emotion instruction. Used by Qwen 3 and Gemini Flash (sent as style instructions); ignored elsewhere. |
| temperature | number, 0–2 | — | Used by Qwen 3, Orpheus, Chatterbox HD, Gemini Flash; ignored elsewhere. |
| top_p | number, 0–1 | — | Qwen 3 only; ignored elsewhere. |
language by model: Qwen 3 → full names (English, Chinese, …; default auto); xAI → ISO 639-1 (en), passed through as given, so send a valid code (default auto); ElevenLabs Turbo → ISO 639-1 (values longer than 5 chars dropped, shorter ones forwarded); MiniMax → full names (sent as a language boost, unchecked); Gemini Flash → full locale strings such as English (US) or Japanese (Japan) (anything else dropped). Kokoro, Inworld, Chatterbox, Orpheus and Gradium ignore it.
Models
Every id below is in the live GET /models?type=tts list. Prices are model_spec.pricing.input.usd = USD per 1M input characters.
| Model ID | Default voice | Formats (default first) | Privacy | Notes |
|---|---|---|---|---|
| tts-kokoro | af_sky | mp3, opus, aac, flac, wav, pcm | private | Multilingual via voice prefix. speed passed through unclamped. |
| tts-qwen3-0-6b / tts-qwen3-1-7b | Vivian | mp3 | private | prompt, temperature, top_p, language. |
| tts-xai-v1 | eve | mp3, wav, pcm | anonymized | 26 voices, ISO language. |
| tts-inworld-1-5-max | Craig | wav | anonymized | Low-latency; all voices are English. |
| tts-chatterbox-hd | Aurora | wav | private | temperature. Voice cloning (zero-shot). |
| tts-orpheus | tara | wav | private | temperature. |
| tts-elevenlabs-turbo-v2-5 | Rachel | mp3 | anonymized | Accepts raw ElevenLabs Voice IDs as voice. |
| tts-minimax-speech-02-hd | WiseWoman | mp3, pcm, flac | anonymized | language. Cloning with this model is not available to regular keys (see below). |
| tts-gemini-3-1-flash | Kore | mp3, opus, wav | anonymized | prompt, temperature, locale language. |
| tts-gradium-v1 | Emma | wav, pcm, opus | anonymized | The voice picks the language; no language param. |
Always inspect GET /models?type=tts before calling: model_spec.voices (authoritative voice list), supported_formats, default_format, supports_custom_voice_id, privacy, pricing, and — on cloning models — voice_cloning. Per-model parameter support (prompt / temperature / top_p / language) is not exposed on /models; use the table above.
A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized models (PRIVATE_TEXT keys are not restricted here).
Voices (from model_spec.voices)
Case-sensitive. Omit voice to get the model's default.
- Kokoro —
_:aAmerican,bBritish,zChinese,fFrench,hHindi,iItalian,jJapanese,pPortuguese,eSpanish;f/mgender. Examples:af_sky,af_bella,af_heart,am_adam,am_michael,bf_emma,bm_george,ff_siwis,jf_alpha,zf_xiaoxiao,ef_dora,pm_alex(54 total). - Qwen 3 —
Vivian,Serena,Ono_Anna,Sohee,Uncle_Fu,Dylan,Eric,Ryan,Aiden - xAI —
eve,ara,rex,sal,leo,altair,atlas,carina,castor,celeste,cosmo,helios,helix,iris,kepler,lumen,luna,lux,naksh,orion,perseus,rigel,sirius,ursa,zagan,zenith - Orpheus —
tara,leah,jess,mia,zoe,leo,dan,zac - Inworld —
Craig,Ashley,Olivia,Sarah,Elizabeth,Priya,Alex,Edward,Theodore,Ronald,Mark,Hades,Luna,Pixie - Chatterbox —
Aurora,Britney,Siobhan,Vicky,Blade,Carl,Cliff,Richard,Rico - ElevenLabs Turbo —
Rachel,Aria,Sarah,Laura,Charlotte,Alice,Matilda,Jessica,Lily,Roger,Charlie,George,Callum,River,Liam,Will,Eric,Chris,Brian,Daniel,Bill— or any ElevenLabs Voice ID - MiniMax —
WiseWoman,FriendlyPerson,InspirationalGirl,CalmWoman,LivelyGirl,LovelyGirl,SweetGirl,ExuberantGirl,DeepVoiceMan,CasualGuy,PatientMan,YoungKnight,DeterminedMan,ImposingManner,ElegantMan - Gemini Flash —
Achernar,Achird,Algenib,Algieba,Alnilam,Aoede,Autonoe,Callirrhoe,Charon,Despina,Enceladus,Erinome,Fenrir,Gacrux,Iapetus,Kore,Laomedeia,Leda,Orus,Pulcherrima,Puck,Rasalgethi,Sadachbia,Sadaltager,Schedar,Sulafat,Umbriel,Vindemiatrix,Zephyr,Zubenelgenubi - Gradium — English:
Emma,Kent,Eva,Jack. German:Mia,Maximilian. Spanish:Valentina,Sergio. French:Elise,Leo. Portuguese:Alice,Davi
A voice not in the chosen model's list (and not a valid handle / custom Voice ID where allowed) → 400. An ElevenLabs Voice ID that the provider rejects also → 400 (not charged).
Voice cloning — POST /audio/voices
Clone a voice from an audio sample and get back a handle (vv_…) to pass as voice on /audio/speech. multipart/form-data only; max 25 MB.
curl https://api.venice.ai/api/v1/audio/voices \
-H "Authorization: Bearer $VENICE_API_KEY" \
-F "model=tts-chatterbox-hd" \
-F "file=@sample.wav"
{ "id": "vv_…", "model": "tts-chatterbox-hd" }
| Field | Notes |
|---|---|
| file | The voice sample, multipart field file. Validated by extension/MIME and binary signature. Aim for a clean speech recording of at least 5 s (voice_cloning.min_sample_seconds; advisory — Venice does not measure duration). |
| model | Defaults to tts-chatterbox-hd, the only cloning model open to regular API keys. |
tts-chatterbox-hd advertises its cloning contract on /models as model_spec.voice_cloning:
{ "mode": "zero_shot", "accepted_formats": ["mp3", "wav", "flac", "mp4"], "min_sample_seconds": 5, "retention_days": 7 }
- Zero-shot: no voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles stop working 7 days after creation, regardless of use.
mp4covers M4A. Samples in other containers (or with a non-audio signature) →400before anything is uploaded.- Each successful clone is charged a flat per-clone fee; synthesis is billed separately per character on
/audio/speech.
tts-minimax-speech-02-hd also appears in the model enum in the OpenAPI spec, but cloning with it isn't open to regular keys, which get 403 "Voice cloning … is not available on your account". Its model spec on /models carries no voice_cloning object — use that as the signal.
A handle is bound to the model that created it. Pass it with a model that has no cloning support → 400; pairing it with a different cloning model fails.
Streaming
{
"model": "tts-xai-v1",
"voice": "eve",
"input": "Hello, this is a long document to narrate. ...",
"streaming": true,
"response_format": "mp3"
}
With streaming: true, the body is a chunked (Transfer-Encoding: chunked) audio stream — decode as it arrives. pcm (where supported: Kokoro, xAI, MiniMax, Gradium) is convenient for raw Web Audio playback. If the stream fails before headers are sent you get 500 {"error":"Stream error"}; after that the connection just ends.
OpenAI SDK
import OpenAI from 'openai'
import fs from 'node:fs/promises'
const client = new OpenAI({
apiKey: process.env.VENICE_API_KEY,
baseURL: 'https://api.venice.ai/api/v1',
})
const mp3 = await client.audio.speech.create({
model: 'tts-xai-v1',
voice: 'eve',
input: 'Hello from Venice.',
response_format: 'mp3',
})
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer()))
Style / emotion (Qwen 3, Gemini Flash)
{
"model": "tts-qwen3-1-7b",
"voice": "Vivian",
"input": "We did it!",
"prompt": "Excited and energetic.",
"temperature": 0.9,
"top_p": 0.95
}
tts-gemini-3-1-flash also takes prompt (style instructions) and temperature, but not top_p. For other families, delivery comes from the voice choice itself (e.g. Inworld Hades vs Pixie); prompt / temperature / top_p are silently ignored.
Errors
| Code | Meaning |
|---|---|
| 400 | Missing model, schema error (strict body, input > 4096 / empty / unspeakable), voice not valid for the model, unsupported response_format, bad cloning sample (/audio/voices), handle paired with a non-cloning model. |
| 401 | Authentication failed. |
| 402 | Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or "API key USD|DIEM spend limit exceeded…" when the key's own cap is hit (no code field); x402 → PAYMENT_REQUIRED with top-up info. |
| 403 | A PRIVATE_ONLY key calling an anonymized model, region restriction, or a cloning model not open to your account on /audio/voices. |
| 404 | Unknown model. |
| 413 | /audio/voices sample over 25 MB. |
| 429 | Rate limited. |
| 500 | Inference failure / stream error. |
| 502 | Temporary upstream TTS failure — {"error":"Speech synthesis failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff. |
| 503 | Model temporarily offline — retry with jitter. |
See venice-errors for body shapes and retry strategy.
Gotchas
- Always send
model— the documented default never applies. inputhard cap is 4096 chars. For long content, split on sentence boundaries and concatenate audio client-side.- Don't assume
mp3: Inworld, Chatterbox, Orpheus and Gradium default towav, and anyresponse_formatoutside a model'ssupported_formatsreturns400. Omitresponse_formator check/modelsfirst. speedis only applied by Kokoro (unclamped), xAI, ElevenLabs Turbo and MiniMax (clamped). Keep0.8–1.3for natural narration.streamingis a Venice-specific field that isn't in the OpenAI SDK's types; pass it as an extra body field, or call the REST endpoint directly and consume the body.- Voice names are case-sensitive (
eve≠Eve,af_sky≠AF_SKY). Noteleo(xAI/Orpheus) vsLeo(Gradium, French). - Chatterbox cloned handles expire 7 days after creation. Re-clone rather than storing handles long-term.
- Gradium has no
languageparameter — pick the voice for the language you want.