create-video-fal Skill
通过任意 FAL 视频模型(Kling、Seedance、Veo)进行图生视频(或文生视频),必须经由 GooseWorks fal-proxy 路由,以便计费到 Ads agent。模板配方指定模型和参数;image_url 输入必须是公网 URL(编排器通过 MCP media_upload 托管本地帧)。返回结果视频 URL 并下载。用于任何视频广告格式的生成式基础片段。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
create-video-fal
Image-to-video (or text-to-video) via any FAL video model (Kling, Seedance, Veo), ROUTED THROUGH THE GooseWorks fal-proxy so the call bills the Ads agent. The template recipe names the model + params; image_url inputs must be public URLs (the orchestrator hosts local frames with the MCP media_upload). Returns the result video URL and downloads it. Use for the generative base clip of any video-ad format.
Run
gen_video.py --model fal-ai/kling-video/v3/pro/image-to-video --payload '{...}' --out clip.mp4 — bills the agent; host-swaps the FAL queue URLs; downloads the result.
Contract
- Paid calls route through the GooseWorks proxies (bills the Ads agent) via the
bundled media_proxy.py — never a provider SDK's default host.
- The template recipe (DB) supplies the model + params; this capability is generic.
Rejection, physical constraints and cast planning
A provider likeness/policy rejection stops the attempt. Preserve the provider's reason, request id and charged/uncharged/unknown state. Do not resubmit an identical rejected payload. Offer a permitted original character, user-cleared reference, or a supported non-likeness route only when allowed by that provider. A different model is not a policy bypass. Review changed inputs and extra spend through the normal approval flow.
Before generation, write a scene checklist from the brief: each wearable's exact count and body location; which hand holds each object; allowed gestures; object contacts and movement; cast identities and reference ownership. Keep unnecessary hands still, use one simple action per shot, and review the whole generated take against the checklist. A prompt is prevention, not proof: reject extra/missing products, impossible contacts or identity drift.
For multiple characters, compare a shared scene with pinned references, fewer people per shot, and separately generated/composed plates. The first preserves interaction but risks identity drift; separate plates improve control but add composition work and may weaken interaction. Lock an approved reference per person and map who speaks each line. No six-character/two-attempt guarantee is supported. A future paid benchmark must state cast size, attempts, budget, model/settings and pass criteria (identity, speaker, counts, gestures and complete dialogue) and retain every failure.
Model notes
How the video models behave, measured on shipped projects. Each note lives here once; recipes
point here instead of repeating it. Seedance and the talking-creator models keep their notes in
their own atoms.
- Veo 3.1 reads states as stills. "Legs in jeans on a wet curb" gives a near-static clip. Name
actions with verbs (steps off, pivots, taps) when the clip has to move.
- Veo 3.1 keeps the start image's composition (angle, distance, what is cropped) for the whole
clip. The start image is the framing, not just a reference.
- Veo 3.1 takes 4, 6 or 8 seconds only. It paces speech to fill the length, so ask for about 6
rather than 8 and keep the tail short; defects live in the tail. Handheld motion and a music bed
are luck per seed, so plan several seeds for a shot that needs them.
- Veo 3.1 Fast ad-libs words and, from about 4 seconds, the eyes can widen and stare. Ask for 4
seconds for a one-line clip.
- Kling v3 for flat 2D or editorial illustration: at
cfg_scale0.5 or lower, with motion-only
prompts, it adds on-style motion where Seedance and Veo invent photoreal middle states.
- Kling v3 holds one state per clip. Staged "first, then, finally" prompts change only in the
last half-second, and a neutral face drifts to a smile and then to mouthing words. Ask for one
state per clip and build the arc in the edit.
- Kling v3 standard returns 2:3 (784x1176) when asked for 9:16. Check the returned size and give
the start image side margin when Kling is the engine.
- Hard constraints need saying three ways (image and video). A single "no face" holds about
half the time. State it as a positive rule near the top, a negation in the middle and a scope
rule at the end.
- Failures: an NSFW false positive needs the visual trigger words removed; a timeout with no
detail gets one unchanged retry, then a simpler prompt; a rate limit means wait. A policy
refusal is final (above).
- Realism comes from real pixels. Prompted "grainy, shot on a phone" still reads as generated.
When a shot must look filmed, restyle real footage, and narrow what the model invents: a blank
glowing screen generates well, a legible interface does not, so composite the real UI in.
- Change over a long period cannot be generated. A shot defined by change longer than one clip
(screens changing through a work session) needs real footage; the model renders one moment.