create-vo-elevenlabs Skill
通过 ElevenLabs 文字转语音生成配音(VO)片段,必须经由 elevenlabs-proxy 路由,以便计费到 Ads agent。voice id 和脚本文本来自模板配方。用于 VO 驱动的视频广告格式(cgi-app-sizzle、flat-vector-explainer、hypermotion)的口播旁白。绝不直接调用 ElevenLabs——必须经过代理归因。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Superseded: the video kit now does this with the voice-elevenlabs part, version 1.0.1, in the parts folder of this repository. It speaks every scene's line through the private line, keeps character timings as word timings and swaps in confirmed pronunciations the same way. This atom stays, unchanged in behaviour, for skills outside the kit until they move; its scripts still run.
create-vo-elevenlabs
VO via the elevenlabs-proxy (bills the Ads agent). gen_vo.py --text "..." --voice --out vo.mp3.
The template recipe supplies the voice + text; paid call routes through media_proxy.
Brand pronunciations
A brand name or product name the voice gets wrong is spoken from its phonetic
spelling, never from a (pronounced …) note (the voice reads notes aloud):
gen_vo.py --text "Meet Drinkag1." --voice <id> --out vo.mp3 --rules working/brand-rules.json
# or: --say-as "Drinkag1=drink A G one"
Only the SPOKEN text changes. Captions, on-screen text and the approved script
keep the written name. When captions are timed from this VO, map the spoken words
back to the written term.
Read confirmed pronunciation before spending
At the start of every video, read the selected brand's persistent facts, including user-authored pronunciation rules. A previous video's notes or a cached rules file do not establish current pronunciation. For an unknown or ambiguous name, resolve how it is said before the paid voice step. Save the user's confirmed choice in the existing brand-facts store, update an existing rule by its id, then make a separate read and confirm it survived. Never infer successful persistence from a write acknowledgement.
The bundled read_pronunciations.py exports user-authored must rules from a fresh context snapshot (the brand read's payload, its {"result": …} wrapper, or the whole tool result). A rule's text is Pronounce "" as ""; the older Pronunciation: => is also read. It rejects a different brand, conflicting rules and required names that remain unknown. The resulting local rules file is a per-run voice input; the brand store remains the durable source. gen_vo.py supports a free dry run to review the resulting spoken text. Captions retain written names.
Model notes
How the ElevenLabs voices behave, measured on shipped projects. Each note lives here once.
- eleven_v3 timings break around audio tags (
[curious],[pause]): tag characters get
invented times and nearby characters bunch together. For captions or sound sync, re-transcribe
the rendered audio with fal-ai/whisper at word level.
- eleven_v3 length varies 5 to 15% between renders of the same text and settings. Fix the
picture length first and fit the voice to it; if that needs more than about 1.4x tempo, cut tags
and render again.
- eleven_v3 pace is bimodal across takes and can run slow for a whole day. Make six takes and
pick by words per second; before a lip-sync locks a take, check one take's pace on a line whose
usual pace you know.
- eleven_multilingual_v2:
styleis the expressiveness dial (0.05 unless there is a reason);
the text cannot buy pauses (an ellipsis or a break tag changes nothing), so pad silence after
generation; speed is the only rate control (0.85 lands in conversational range, and the
scale is close to linear from 0.80 to 1.00); stability changes delivery, not rate.
Character timings for video
The bundled generator supports timestamped voiceover with per-voice settings.
Use the timestamp option for podcast dialogue and other edits that must follow
speech. It saves an MP3 and a character-alignment JSON beside it, sharing the
same filename stem. The existing audio-only invocation keeps its behavior.
A missing alignment is an error; do not render captions from guessed timings.
See the bundled timestamp tests for the request and output contract.