From cb1423a61ff4a5cdf22779eea1668a6f67dd90b5 Mon Sep 17 00:00:00 2001 From: calesthio Date: Sat, 4 Apr 2026 12:51:47 -0700 Subject: [PATCH] Add audio architecture decision to video-reference-analyst planning The agent must lock the audio approach (single narrator vs. character dialogue vs. both) during Step 3, not defer it to script/compose stage. Proposals now include voice casting with specific voice IDs. --- skills/meta/video-reference-analyst.md | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/skills/meta/video-reference-analyst.md b/skills/meta/video-reference-analyst.md index 8198b01..2c48fe2 100644 --- a/skills/meta/video-reference-analyst.md +++ b/skills/meta/video-reference-analyst.md @@ -132,10 +132,18 @@ hardcode key names or setup URLs. Before proposing, gather what the VideoAnalysisBrief doesn't tell you: 1. "Do you want narration in your version, or visuals-only with music?" -2. "How long should your video be? The reference is [X] seconds." -3. "Is there a specific topic/subject you want, or should I riff on the +2. **If narration: lock the audio architecture now.** Ask: + "How should the story be told? Options: + • **Single narrator** — one voice tells the whole story (like a Pixar short) + • **Character dialogue** — characters speak to each other, no narrator + • **Narrator + character voices** — narrator drives the story, characters + have occasional dialogue lines" + This decision shapes the script, voice casting, and budget. It MUST be + resolved before proposals — do not defer it to the script or compose stage. +3. "How long should your video be? The reference is [X] seconds." +4. "Is there a specific topic/subject you want, or should I riff on the same theme as the reference?" -4. "Any elements from the reference you specifically love or hate?" +5. "Any elements from the reference you specifically love or hate?" Do NOT ask all at once. Lead with the most important gap. If the user's initial message already answers some of these, skip those. @@ -195,7 +203,9 @@ Use this structure for each variant: - Motion: [video gen clips / Remotion spring animations on stills / etc.] **Audio plan:** -- Narration: [yes/no, which TTS provider, voice style] +- Audio architecture: [single narrator / character dialogue / narrator + characters] +- Voice casting: [voice name + ID for each role — narrator, character A, etc.] +- TTS provider: [ElevenLabs / other] - Music: [library track / generated / none] - Sound design: [any special audio needs]