stage-assemble Skill
将已批准的跨模态 EDL(plan.json)确定性地组装成成片——逐段产出内容(edit/compose/generate/provided,分别委派给对应流程),然后用 ffmpeg 分层组装:拼接主轨 → 叠加合成图层 → 统一混入解说(带覆盖检查)→ 烧录字幕 → 校验响度;支持幂等续跑,并在草稿门禁前进行 QA。在计划通过校验并获批后触发;不要用于导入或重新编写计划(stage-plan)。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
stage-assemble
How to execute a validated project/plan.json into one finished file. By the time you are here the plan passed ovs plan validate and the user approved it at gate B. Walk it; do not re-plan. The producers are ovs edit, ovs draft, ovs video/ovs image, ovs speak, and the assembler is ovs edit (or the equivalent MCP tools).
Step 1 — Produce each segment (delegate by source)
Iterate segments in order. For each, produce its produced_path according to source, then write that path + status:"done" back into the segment so a resume never re-produces it:
- edit →
stage-edit:ovs edit trimtheinput_idto[in_sec, out_sec]→project/cuts/.mp4. - compose →
stage-compose: build a small visual-only manifest-owned composition forspec.kind(title card, lower-third, stat card, captions) underproject/compositions//→ runovs draft project/compositions/ --out project/parts/.mp4 --quality draft --report project/reports/-compose-report.json. This keeps compose segments on the same manifest/source/check/video-QA path as standalone COMPOSE while still letting the assembler own narration and loudness. - generate →
stage-generate(+stage-consistencyfor recurring characters): only AFTER gate C.ovs video/ovs image→project/assets/.mp4. Foroperation:"edit", pass the exact original reference video and obey top-levelreferencesplusedit_strategy; never widen it into regeneration. A failed/unknown paid attempt is not an automatic retry. Preserve completed siblings and require a new output path for any later authorized attempt. - provided → use
spec.asset_idas-is (probe it first; conform aspect/fps if needed).
Billable generate segments must not run before gate C has confirmed the count from cost_estimate. Produce cheap/free segments (edit, compose, provided) freely.
Step 2 — Assemble in ffmpeg tiers (the default path)
Assemble deterministically, bottom-up. This tiered order is the default; it is predictable and cheap, and keeps each clip's real audio intact:
- Primary track —
ovs edit concatthe primary-layerproduced_paths inorder→project/render/primary.mp4. Conform aspect/fps on the way in if sources differ; read the returned conformance report and verify it still matches the approved canvas. - Overlays / bg — Do not place a full-frame opaque composed video over source footage: it erases the base. Use genuinely transparent or bounded overlays, or amend the plan to make it a primary beat. Do not bypass the opaque-overlay refusal with resizing tricks. for each overlay/bg segment,
ovs edit overlayits part onto the primary over the window of the segment named inover(title cards, lower-thirds, logos). Composed layers are VISUAL-ONLY — they must not carry their own narration audio. This includes a compose segment that IS the primary track (a full-video composition): render it SILENT — do not put a narrationin itsindex.html. The assembler owns narration (tier 3), so a composition that bakes it in would mean narration is added TWICE (the "two voices" defect). - Narration — added EXACTLY ONCE, here. If active, require the Gate-B-signed
tracks.narration.synthesisprofile. Runovs narration fitfor each timed line beforeovs speak; shorten over-budget text in the plan without changing the approved meaning. Synthesize with the exact signed voice/model/format/speed, probe the result, rerun measured fit, and write each line'sproduced_path. If measured timing misses, revise once using the suggested unit budget rather than repeatedly billing or forcing speed. Add all produced lines in ONEovs edit mixcall at theirstart_sec. The default existing-audio rejection catches compose segments that accidentally baked narration; re-render those SILENT. ReadvoicedRatio,interiorGaps,maxOverlapSecand status: the last line reaching the end does not mean the whole track carries speech. Fix unintended dead air and collisions, and disclose any intentional silent tail at Gate D. Preserve a generated talking head's built-in lip-synced audio instead of adding a second voice. - Music — add
tracks.musicducked under narration by the planned amount. - Captions — turn
tracks.captions.lines({text, start_sec, target_sec}) into a.srt, thenovs edit burnsubs. Avoid duplicating caption text already visibly present in that scene. Captions are DATA in the plan — burned ONLY here at assemble — so a later typo fix reuses the clean pre-caption primary/generated source and burns the revised file once. Never burn onto the already-captioned output and never regenerate an unchanged provider clip. Ifburnsubsfails because the runtime ffmpeg lacks subtitle filter support, stop and report that blocker; do not hand-write a fallback ffmpeg graph, PNG subtitle overlay, or drawtext pipeline. - Loudness — run
ovs edit normalize-loudness project/render/draft.mp4 --out project/render/video.mp4. It normalizes to thevideo-craft§7 targets (~−14 LUFS integrated, true-peak ≤ ~−1 dBTP) and returns measured loudness; useovs edit loudnessonly for diagnosis without writing an output.
Apply the plan's style_kit for cohesion: composed layers (titles/captions/cards) use its palette + fonts. A single lut graded across all clips is what unifies tonally mixed sources — until a grade op is available, keep mixed sources close at capture/trim and lean on the shared palette + consistent captions for cohesion rather than promising a uniform grade.
Output project/render/video.mp4 as the deliverable; project/render/draft.mp4 is the pre-normalized intermediate.
Director judgment (end-to-end assembly)
The craft of making mixed sources feel like one video, on top of the shared craft (video-craft). The seams between footage / generated / composed are where multi-source assembly falls apart — engineer continuity across them:
- One look across every source. Apply the
style_kitso a cut from real footage → a generated shot → a composed card does not read as three videos: one type system + palette on every composed layer, one caption style throughout, matched aspect / fps, tonal proximity (a shared LUT is the unifier when available;video-craft§4). - Audio is the through-line that hides the visual seam. One narration voice; a continuous music bed UNDER the cuts (do not restart it per segment); duck consistently (
video-craft§7). The ear's continuity carries the eye across a source change — a reveal may drop music, but the bed bridges the cut. - Rhythm over a mixed cut. Alternate motion vs. static and source types for momentum — do not stack three composed cards or three talking-head shots in a row (that is the repetition / slideshow smell,
video-craft§3, §12). Vary holds. - Cut on a content change, not just plan order. A hard cut on a beat / word change is invisible and professional; a crossfade signals a gentle topic shift (
video-craft§5). - Don't bury the hero. On a
source_ledpiece, composed lower-thirds and captions FRAME the footage — they never cover its subject / face (video-craft§6). - Apply the editing cut craft ACROSS the seams. The cut mechanics live in
stage-edit→ "Cut craft" (best sub-window, ≤ 4 transitions, L/J-cut sound bridges, handles / no freeze-frame, adjacent-diversity, a reason per cut) — apply them at every junction between sources, since the footage → generated → composed seams are exactly where a mixed cut betrays itself.
Step 3 — Idempotent resume
The plan is the checkpoint. On a re-run, skip any segment already status:"done" with a present produced_path, and skip assembly tiers whose output already exists and is newer than its inputs. Never re-run a billable generate segment that is already produced.
For a partial child failure or later revision, reset only the affected child and assembly tiers derived from it. Preserve every unaffected completed child and its evidence. When the child becomes valid, rebuild the complete parent draft and run parent QA in the same workflow; do not stop at a child-only recovery artifact.
Step 4 — QA report, then gate D
Before showing the draft, run the QA pass and write project/render_report.json with these sections:
- technical_probe —
ovs edit probethe draft/final (real duration / resolution / fps / audio present); confirm it matches the plan's aspect + total. - promise_preservation —
ovs plan promise-check project/plan.json --probe-produced. At gate D this probes each primary segment'sproduced_pathand computes the REAL primary-track motion ratio vs.motion_min_ratioplus thesource_requiredinvariant; missing/unreadable produced media or a fail means "slideshow / promise broken" — do not deliver. Send it back (below). Do not eyeball this; let the numbers decide. - visual_spotcheck — extract ~4 frames across the draft (
ovs edit extract-frame) and read them for upside-down / garbled-caption / empty / wrong-product frames. Read them yourself if you are multimodal; if you cannot see images, record the spot-check asunverifiedand proceed — do not invent what the frames show. - audio_spotcheck — the
normalize-loudnessmeasured loudness numbers + the narration coverage result from step 2 (uncovered tail / overshoot / silent lead-in). - transcript_comparison (when there is narration) — optionally
ovs transcribethe draft and confirm the spoken words match the planned narration lines.
Each section carries pass / warn / fail + a one-line reason. Then present the draft + headline findings at Gate D and resolve the user's decision through ovs gate transition.
On approve → finalize project/render/video.mp4 (loudness / captions only; never re-synthesize a talking-head voice). On revise → redo only the affected segment(s) and re-assemble.
Send-back (self-correction on a QA fail)
A QA fail does not go to the user as "here's a broken video". Diagnose which segment(s) caused it and redo ONLY those, then re-assemble and re-run QA:
- promise_preservation fail (slideshow) → the static composed segments are too long / the motion segments too short. Rebalance segment durations or convert a static beat to footage, re-assemble.
- visual_spotcheck fail (bad frame) → re-produce that one segment (re-trim / re-compose / re-generate), not the whole video.
- audio fail (uncovered tail) → re-time or extend the narration / trim the tail.
Bound repetition, not recovery: allow at most 2 send-back rounds for the same failing check and unchanged strategy. Then preserve the current artifact and show concrete user directions through gate-control before starting another cycle. Create no technical confirmation. Only a required signed-plan change or new billable attempt returns through its normal authorization boundary.
Rules
- Walk the approved plan; if assembly reveals the plan is wrong, surface it and re-gate — do not silently re-plan.
- Write
produced_path+statusback per segment as you go (resumability + the QA pass depend on it). - One output file is the deliverable;
cuts/andparts/are intermediates. - Narration is added exactly ONCE — in the mix tier, never baked into a compose render. Compose segments (including a full-video composition used as the primary track) render SILENT (no narration
); the assembler mixes narration viaovs edit mixwithsegmentsplaced per line. The mix's default--on-existing-audio rejectenforces this — a "base already has audio" mix rejection is the signal a segment wrongly baked audio in; re-render it silent, then re-mix. - No ad-hoc ffmpeg fallbacks for captions. Caption burn-in is a low-freedom operation owned by
ovs edit burnsubs; a failed burnsubs call is a tool/runtime blocker, not permission to invent a custom subtitles/drawtext/PNG-overlay command.
Boundary / non-goals
This skill assembles an already-approved plan. It does not ingest or decide the plan (stage-plan), and it delegates the actual production of each segment to the compose / generate / edit / consistency skills rather than re-deriving their craft here.
Before showing any plan-backed final, regardless of assembly route, run ovs plan promise-check project/plan.json --probe-produced --video project/render/video.mp4. Repair duration/aspect mismatch, overlapping/truncated narration and unverifiable lines. Explain caption warnings from the actual container/sidecar check; burned captions need visual evidence. Planned duration and a successful mix alone do not prove delivery.
For narration, preserve signed line windows; target_sec is duration. Measure each produced file and voice span. Keep intentional silence, shorten overlong text within authorized scope, and never slide later lines to hide an overlap. A failed/unknown provider request does not authorize automatic repeat billing.