video-gen: adopt Seedance 2.0 as preferred premium default
Seedance 2.0 is now routed as the top-ranked video generator whenever a premium gateway is available. Touches the tool layer, scoring engine, cinematic pipeline, and both skill layers so discovery works from every entry point. - tools/video/seedance_video: BETA stability, quality_score=0.95, add reference_to_video operation plus 9 img + 3 vid + 3 audio ceilings, fix pre-existing upload_image_fal import - tools/base_tool: surface optional quality_score / success_rate / latency fields in get_info so the scorer can read them - lib/scoring: fix reliability enum-vs-string bug that was pinning every available tool to 0.0, switch to overlap coefficient so rich best_for descriptions aren't penalized, add premium-cinematic feature bonus - pipeline_defs/cinematic + cinematic asset-director: add pixabay_music and freesound_music, restore pixabay-first music default - cinematic compose-director: mandatory Remotion preflight at stage entry - New Layer 3 .agents/skills/seedance-2-0/SKILL.md (8-part prompt structure, multi-shot, lip-sync, reference-to-video, provider landscape) - New Layer 2 skills/creative/prompting/seedance-prompting.md - Update ai-video-gen, video-gen-prompting, AGENT_GUIDE, INDEX to flag Seedance 2.0 as the preferred premium default and make the skill discoverable from every routing path
This commit is contained in:
@@ -92,6 +92,7 @@ Key capability families to look for in the output:
|
||||
| Data Visualization | `creative/data-visualization.md` | Chart type selection, animation, label placement | `d3-viz`, `remotion-best-practices` |
|
||||
| Video Stitching | `creative/video-stitching.md` | Multi-clip assembly, AI clip chaining, spatial composition | `ffmpeg`, `video_toolkit` |
|
||||
| Video Gen Prompting | `creative/video-gen-prompting.md` | Universal video generation prompt vocabulary | `ai-video-gen`, `ltx2`, `create-video` |
|
||||
| ↳ Seedance Prompting | `creative/prompting/seedance-prompting.md` | **Preferred premium default.** Seedance 2.0 8-component structure, multi-shot, lip-sync, reference-to-video | `seedance-2-0`, `ai-video-gen` |
|
||||
| ↳ Grok Prompting | `creative/prompting/grok-prompting.md` | Grok image/video prompting, edit flows, reference-image video | `grok-media` |
|
||||
| ↳ Sora Prompting | `creative/prompting/sora-prompting.md` | Sora 2 structured template, advanced fields | `ai-video-gen` |
|
||||
| ↳ VEO Prompting | `creative/prompting/veo-prompting.md` | VEO 3.1 14-component structure, art movements | `ai-video-gen` |
|
||||
@@ -306,4 +307,5 @@ Claude Code accesses them via symlinks in `.claude/skills/`.
|
||||
| **Animation** | `framer-motion`, `lottie-bodymovin` | `pproenca/dot-skills`, `dylantarre/animation-principles` |
|
||||
| **Design** | `tailwind-design-system`, `web-design-guidelines`, `vercel-react-best-practices`, `vercel-composition-patterns` | `wshobson/agents`, `vercel-labs/agent-skills` |
|
||||
| **AI Video (HeyGen)** | `heygen`, `avatar-video`, `create-video`, `faceswap`, `ai-video-gen`, `video-download`, `video-edit`, `video-translate`, `video-understand`, `visual-style` | `heygen-com/skills` |
|
||||
| **AI Video (Premium)** | `seedance-2-0` — preferred premium default (cinematic, trailer, multi-shot, lip-sync, synced audio); accessed via `seedance_video` (fal.ai) or `heygen_video` Avatar Shots | Local OpenMontage skill |
|
||||
| **Infrastructure** | `acestep`, `ltx2`, `playwright-recording` | `digitalsamba/claude-code-video-toolkit` |
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
# Seedance 2.0 — Prompting Guide
|
||||
|
||||
> Layer 3 authority: `.agents/skills/seedance-2-0/SKILL.md`
|
||||
> For universal vocabulary, see: `skills/creative/video-gen-prompting.md`
|
||||
|
||||
## When to pick Seedance 2.0
|
||||
|
||||
Seedance 2.0 (ByteDance Seed team, released Feb 2026) is OpenMontage's **preferred premium default for cinematic, trailer, teaser, hype-edit, and motion-led clip work** whenever a paid gateway is configured (`FAL_KEY` via `seedance_video`, or HeyGen Video Agent / Avatar Shots). It is the only model in the fleet that delivers all of:
|
||||
|
||||
- single-pass native synchronized audio (speech + SFX + ambience together, not post-sync),
|
||||
- multi-shot generation inside a single prompt,
|
||||
- director-level camera control,
|
||||
- lip-sync from quoted dialogue,
|
||||
- reference-conditioned generation with up to 9 images + 3 video clips + 3 audio clips,
|
||||
- consistent character identity across shots.
|
||||
|
||||
Elo 1269 on Artificial Analysis as of release — ahead of Veo 3, Sora 2, Runway Gen-4.5.
|
||||
|
||||
Switch off Seedance 2.0 only when there is a real reason: strict budget (use the `fast` variant or LTX), explicit user preference (VEO/Sora/Kling), or a stylistic fit another model does better (VEO for photoreal landscape, Kling for anime).
|
||||
|
||||
## Seedance 2.0 8-Component Prompt Structure
|
||||
|
||||
Seedance is unusually literal about camera language, multi-shot cuts, and quoted dialogue. Use this structure — include what matters, omit what doesn't:
|
||||
|
||||
1. **Shot / framing** — wide establishing, medium, close-up, Dutch angle, etc.
|
||||
2. **Camera movement** — static, slow push-in, aerial, handheld, arc, dolly zoom
|
||||
3. **Subject description** — the physical detail that must persist across shots (identity anchor)
|
||||
4. **Action beats** — one beat per sentence, use `→` or explicit `Shot 1 / Shot 2` for multi-shot
|
||||
5. **Setting / environment** — location, era, weather, time of day
|
||||
6. **Lighting / palette** — one lighting idea, pick and commit
|
||||
7. **Style / grade / era** — "anamorphic lens, teal-orange grade, 35mm film grain"
|
||||
8. **Audio** — ambient, diegetic, music direction (textural only), quoted dialogue for lip-sync
|
||||
|
||||
## Seedance-specific strengths
|
||||
|
||||
| Capability | How to invoke it |
|
||||
|---|---|
|
||||
| **Native synced audio** | Describe the soundscape in the prompt. Leave `generate_audio=true`. |
|
||||
| **Multi-shot in one generation** | Use `Shot 1 (...)`, `Shot 2 (...)` etc. Keep subject description consistent across shots. |
|
||||
| **Director-level camera** | Use unambiguous terms: `slow dolly-in`, `arc shot`, `Dutch tilt`, `aerial push-in`, `handheld with micro-shake` |
|
||||
| **Lip-sync from quoted dialogue** | `Character says: "line."` — each line ≤ ~6 words on fast cuts |
|
||||
| **Reference-to-video** | Use the `reference-to-video` endpoint; name each asset in the prompt (`Reference 1: hero character — ...`) |
|
||||
| **Character identity consistency** | Describe the same physical details in every shot — Seedance uses those as the identity anchor |
|
||||
|
||||
## Multi-shot pattern
|
||||
|
||||
Seedance honors explicit shot lists:
|
||||
|
||||
```
|
||||
Shot 1 (wide aerial establishing, slow push-in):
|
||||
Snow-covered Air Temple at dawn, spires catching first orange light.
|
||||
Wind lifting prayer flags.
|
||||
|
||||
Shot 2 (medium, low angle, handheld):
|
||||
Aang — bald, blue arrow tattoo, orange robes — plants his staff on stone.
|
||||
He squints into the rising sun.
|
||||
|
||||
Shot 3 (extreme close-up, rack focus):
|
||||
Rack focus from the glowing arrow tattoo on his forehead to the distant peaks.
|
||||
Aang says: "It's time."
|
||||
|
||||
Style: anamorphic lens, teal-orange cinematic grade, 35mm film grain.
|
||||
Audio: rising orchestral swell with low taiko pulse, wind, distant wingbeats.
|
||||
```
|
||||
|
||||
## Lip-sync pattern
|
||||
|
||||
```
|
||||
Aang says: "I won't run anymore."
|
||||
Sokka, half a step behind, replies: "Then we fight."
|
||||
```
|
||||
|
||||
- Use `Character says: "..."` / `Character replies: "..."` exactly — mouth shapes key off the quoted strings.
|
||||
- Keep lines short (≤ 6 words on fast-cut shots) to avoid drift.
|
||||
- For a single-speaker monologue, keep the camera close and static on the speaker's shot.
|
||||
|
||||
## Parameter cheat sheet
|
||||
|
||||
| Parameter | Guidance |
|
||||
|---|---|
|
||||
| `duration` | `5`–`8` s hero, `10`–`12` s multi-shot scenes, `4` s inserts. `auto` when unsure. |
|
||||
| `aspect_ratio` | `21:9` trailers, `16:9` broadcast, `9:16` Reels/Shorts/TikTok |
|
||||
| `resolution` | `720p` default. `480p` for cost-capped previews only. |
|
||||
| `generate_audio` | Keep `true` — sync audio is the moat. Strip in compose if unused. |
|
||||
| `model_variant` | `standard` for hero + multi-shot + camera-heavy. `fast` for b-roll, previews, latency-capped jobs. |
|
||||
| `seed` | Lock once a shot composition reads; iterate variants with the same seed. |
|
||||
|
||||
## Iteration strategy
|
||||
|
||||
1. **Block out shape** — `duration=5`, `fast`, one shot. Confirm composition.
|
||||
2. **Lock the seed** — record it in the per-clip README.
|
||||
3. **Upgrade to `standard`** — same seed, tighten camera + lighting language.
|
||||
4. **Extend or multi-shot** — only after the single-shot version is clean.
|
||||
5. **Promote to final** — write the prompt, seed, variant, and duration into the asset manifest so compose can re-render consistent retakes.
|
||||
|
||||
## What to avoid
|
||||
|
||||
| Don't | Why |
|
||||
|---|---|
|
||||
| Four-plus simultaneous actions in one shot | Motion coherence collapses. Split to multi-shot. |
|
||||
| Readable text / logos inside the clip | Text rendering is unreliable. Handle text in Remotion overlay. |
|
||||
| Conflicting lighting (`bright noon` + `neon night`) | Model picks one and ignores the other. |
|
||||
| Long dialogue on fast-cut shots | Lip-sync drifts. |
|
||||
| `fast` variant for slow-mo, multi-shot, or complex camera | Routinely misses on first try. Route to `standard`. |
|
||||
| Request a full multi-instrument score from Seedance | Keep audio direction textural; real scoring belongs in `music` / `pixabay_music` / `elevenlabs` and mixes in compose. |
|
||||
| Bypass `video_selector` without a reason | Loses scoring, fallback, and cost handling. |
|
||||
|
||||
## Integration notes
|
||||
|
||||
- **Cinematic pipeline:** Seedance 2.0 is the default. 21:9, multi-shot for montage, reference-to-video when the brief has a visual bible.
|
||||
- **Animated explainer:** Use Seedance 2.0 only for establishing / mood / cold-open clips — core motion graphics stay in Remotion.
|
||||
- **Screen demo / podcast / clip factory:** Not the right default. Only for stylized cold-opens.
|
||||
- **Cost check:** `standard` at 10 s ≈ $3.03 / clip on fal.ai. `fast` at 5 s ≈ $1.21. Budget in the proposal stage.
|
||||
|
||||
## Example — Airbender trailer hero beat (60 s total trailer, this is shot 3 of 7)
|
||||
|
||||
```
|
||||
Shot 1 (wide aerial, slow push-in, 3s):
|
||||
Snow-covered Air Temple at dawn, spires catching orange light,
|
||||
prayer flags lifting in wind.
|
||||
|
||||
Shot 2 (low angle medium, handheld, 3s):
|
||||
Aang — bald, blue arrow tattoo on forehead, orange and yellow robes —
|
||||
plants his staff on weathered stone, squints into the rising sun.
|
||||
|
||||
Shot 3 (extreme close-up, rack focus, 3s):
|
||||
Rack focus from the glowing arrow on his forehead to distant peaks.
|
||||
Aang says: "It's time."
|
||||
|
||||
Lighting: cold blue ambient with warm break on the horizon,
|
||||
rim light from rising sun.
|
||||
Style: anamorphic 2.39:1, teal-orange cinematic grade, 35mm film grain,
|
||||
halation on speculars.
|
||||
Audio: low taiko drums rising to orchestral swell on Shot 3,
|
||||
wind through temple, distant wingbeats, leather staff-grip creak.
|
||||
```
|
||||
|
||||
Parameters: `duration=10`, `aspect_ratio=21:9`, `resolution=720p`, `model_variant=standard`, `generate_audio=true`, seed locked after shot 2.
|
||||
@@ -2,9 +2,11 @@
|
||||
|
||||
## When to Use
|
||||
|
||||
When writing prompts for the video generation family (`video_selector`, `heygen_video`,
|
||||
`wan_video`, `hunyuan_video`, `ltx_video_local`, `ltx_video_modal`, `cogvideo_video`).
|
||||
This skill covers the universal prompt vocabulary that works across all video generation models.
|
||||
When writing prompts for the video generation family (`video_selector`, `seedance_video`,
|
||||
`heygen_video`, `wan_video`, `hunyuan_video`, `ltx_video_local`, `ltx_video_modal`,
|
||||
`cogvideo_video`). This skill covers the universal prompt vocabulary that works across all
|
||||
video generation models. For the **preferred premium default**, see the Seedance 2.0 row
|
||||
in the table below.
|
||||
|
||||
For model-specific tips, see the linked guides below.
|
||||
|
||||
@@ -12,6 +14,7 @@ For model-specific tips, see the linked guides below.
|
||||
|
||||
| Model | Guide | Key Insight |
|
||||
|-------|-------|-------------|
|
||||
| **Seedance 2.0 (standard / fast)** | `creative/prompting/seedance-prompting.md` + Layer 3 `.agents/skills/seedance-2-0/` | **Preferred premium default** when `FAL_KEY` or HeyGen is configured. Single-pass synced audio, multi-shot generation, director-level camera, lip-sync from quoted dialogue, reference-to-video (9 img + 3 vid + 3 audio). Elo 1269 (#1 on Artificial Analysis). |
|
||||
| **Sora 2 / Sora 2 Pro** | [OpenAI Sora 2 Cookbook](https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide) | Richest structured template. Advanced fields: lenses, filtration, grade, diegetic sound, wardrobe, finishing. |
|
||||
| **VEO 3.1 / VEO 3** | [Vertex AI Prompt Guide](https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide) | Best vocabulary reference tables. 14-component prompt structure. |
|
||||
| **Grok Imagine Video** | `creative/prompting/grok-prompting.md` | Best when prompts need reference-image placeholders like `<IMAGE_1>` and identity/product carryover. |
|
||||
|
||||
@@ -27,7 +27,7 @@ Before authoring title cards, name plates, or SVG overlays, read **`skills/meta/
|
||||
|-------|----------|---------|
|
||||
| Schema | `schemas/artifacts/asset_manifest.schema.json` | Artifact validation |
|
||||
| Prior artifacts | `state.artifacts["scene_plan"]["scene_plan"]`, `state.artifacts["script"]["script"]`, `state.artifacts["proposal"]["proposal_packet"]` | Scene intent and beat plan |
|
||||
| Tools | `subtitle_gen`, `audio_enhance`, `image_selector`, `video_selector`, `music_gen` — selectors auto-discover all available providers from the registry | Optional support asset creation |
|
||||
| Tools | `subtitle_gen`, `audio_enhance`, `image_selector`, `video_selector`, `pixabay_music` (free, default), `freesound_music` (free), `music_gen` (ElevenLabs, paid) — selectors auto-discover all available providers from the registry. **Default to `pixabay_music` before reaching for `music_gen`.** | Optional support asset creation |
|
||||
| Playbook | Active style playbook | Brand and typography consistency |
|
||||
|
||||
## Process
|
||||
@@ -54,7 +54,7 @@ If `proposal_packet.metadata.motion_required = true`, actual moving footage or g
|
||||
Before batch-generating support assets, produce one sample of each expensive generated type and show the user:
|
||||
|
||||
1. **Generated insert sample** (if using `image_selector` or `video_selector`): Generate one representative visual. Confirm it complements the source footage before batching.
|
||||
2. **Music sample** (if using `music_gen`): Generate a short clip. Confirm mood and energy match the beat plan.
|
||||
2. **Music sample** (try `pixabay_music` first — free, searchable by mood/BPM; fall back to `freesound_music` for cues and ambience; only reach for `music_gen` when the search tools miss the brief): sample or retrieve a short clip. Confirm mood and energy match the beat plan.
|
||||
|
||||
If `motion_required = true`, the representative visual must be a video clip sample, not a still image sample.
|
||||
|
||||
|
||||
@@ -24,6 +24,20 @@ If the approved brief or scene plan makes motion a hard requirement, verify that
|
||||
- Do not convert the piece into an animatic unless the user explicitly approves that downgrade.
|
||||
- If the render engine changes materially, tell the user before rendering and explain why.
|
||||
|
||||
**Mandatory Remotion preflight (run before every render when the scene plan includes any Remotion scene type — title cards, stat cards, anime/hero_title, end-tag, overlays):**
|
||||
|
||||
```bash
|
||||
python -c "
|
||||
from tools.tool_registry import registry
|
||||
registry.discover()
|
||||
info = registry.get('video_compose').get_info()
|
||||
print('Render engines:', info.get('render_engines'))
|
||||
print('Remotion note:', info.get('remotion_note'))
|
||||
"
|
||||
```
|
||||
|
||||
If Remotion is not in the available render engines, stop and report to the user per the Decision Communication Contract. Do not substitute a reduced-fidelity render path without approval.
|
||||
|
||||
### 1. Use Frame Treatment Deliberately
|
||||
|
||||
Only use letterbox, 24fps intent, or heavy grading if they help the piece. Do not apply them because the pipeline name says cinematic.
|
||||
|
||||
Reference in New Issue
Block a user