Fix Remotion-first rendering docs and post-render verification gaps
Compose-director had contradictory instructions: Step 2 described Remotion captions/audio, but Steps 5/5b gave detailed FFmpeg code that agents followed instead. This caused three failures in production: FFmpeg subtitles instead of Remotion CaptionOverlay, missing audio (mixed externally but never embedded in Remotion props), and skipped audio verification in post-render review. Changes: - compose-director: Remotion is now DEFAULT for audio, captions, text overlays; FFmpeg is labeled FALLBACK only. Post-render review has mandatory ffprobe gate and audio transcription with explicit stop conditions. - remotion.md: routing table updated (captions/audio → Remotion), added universal Post-Render Verification Protocol for all pipelines (only 2/10 had one). - scene-director, asset-director: added pitfall for AI-generated text in CTA screens — must use Remotion text_card for verbatim text. - image-provider-usage: added Recraft V4 caveat (style param causes 422 on fal.ai). - recraft_image.py: documented the style parameter 422 issue inline.
This commit is contained in:
+37
-1
@@ -30,8 +30,10 @@ simple standalone operations that don't benefit from React rendering.
|
||||
| Video-only cuts with transitions | **Remotion** | Native `<OffthreadVideo>` + transitions |
|
||||
| Animated diagrams/text cards | **Remotion** | Frame-by-frame control |
|
||||
| Data-driven batch videos | **Remotion** | Zod props + parametric renders |
|
||||
| Word-level captions (in composition) | **Remotion** | CaptionOverlay with word highlight — superior to SRT |
|
||||
| Audio embedding (narration + music) | **Remotion** | Native `<Audio>` components with volume/fade |
|
||||
| Simple trim, concat (no composition) | FFmpeg | Instant, no Node dependency |
|
||||
| Subtitle burn-in (standalone) | FFmpeg | Proven, fast |
|
||||
| Subtitle burn-in (standalone, post-hoc) | FFmpeg | Only for adding subs to an already-rendered video without re-rendering |
|
||||
| Face enhance, color grade | FFmpeg | Filter-based, deterministic |
|
||||
| Remotion unavailable | FFmpeg | Automatic fallback |
|
||||
|
||||
@@ -328,13 +330,47 @@ Remotion renders are CPU-intensive but $0 API cost. Track via cost_tracker:
|
||||
- **Node.js 18+ required** — listed as optional in minimum system, required in recommended.
|
||||
- **Render in series, not parallel** — unless the machine has enough RAM. Each render spawns a Chromium instance.
|
||||
|
||||
## Post-Render Verification Protocol (ALL pipelines)
|
||||
|
||||
**Every Remotion render MUST be verified before presenting to the user.** This protocol applies
|
||||
to ALL pipelines, not just explainer. Pipeline-specific compose-directors may extend it but
|
||||
must not skip any step.
|
||||
|
||||
**Step 1: Probe the output file (GATE — blocks all other steps):**
|
||||
```bash
|
||||
ffprobe -v quiet -print_format json -show_format -show_streams rendered_video.mp4
|
||||
```
|
||||
Verify ALL of:
|
||||
- [ ] Video stream exists with correct resolution and FPS
|
||||
- [ ] **Audio stream exists** — if missing, STOP immediately, fix audio config, re-render
|
||||
- [ ] Duration within ±5% of target
|
||||
- [ ] File size is reasonable (not 0 bytes, not suspiciously small)
|
||||
|
||||
**If audio stream is missing, do NOT proceed.** This means narration/music were not embedded.
|
||||
The most common cause: audio sources were mixed externally but never passed in the Remotion
|
||||
`audio` prop. Fix: add `audio.narration` and `audio.music` to composition props and re-render.
|
||||
|
||||
**Step 2: Extract review frames** at scene midpoints and visually inspect each one.
|
||||
|
||||
**Step 3: Transcribe the rendered video's audio** using WhisperX/transcriber tool.
|
||||
- If 0 words returned → audio is silent despite stream existing → investigate
|
||||
- If word count < 80% of script → audio is cut off → investigate
|
||||
- Compare last transcribed word to last scripted word
|
||||
|
||||
**Step 4: Present structured review** to user with file stats, audio verification results,
|
||||
visual findings, and caption status before declaring the video complete.
|
||||
|
||||
## Quality Checklist
|
||||
|
||||
- [ ] Composition duration matches sum of scene durations minus transition overlaps
|
||||
- [ ] All `staticFile()` references resolve to existing assets
|
||||
- [ ] Transitions don't cut off content (account for overlap in timing)
|
||||
- [ ] **Audio stream present in rendered output** (ffprobe confirms codec_type: "audio")
|
||||
- [ ] **Narration words verified via transcription** (not just assumed from props)
|
||||
- [ ] Audio layers are in sync with visual scenes
|
||||
- [ ] Captions/subtitles rendering correctly (Remotion CaptionOverlay preferred over FFmpeg SRT)
|
||||
- [ ] Theme colors match the active style playbook
|
||||
- [ ] Output resolution and FPS match the target media profile
|
||||
- [ ] Render completes without Chromium timeout errors
|
||||
- [ ] Final output plays correctly on target platform
|
||||
- [ ] Text-bearing scenes (CTA, titles) use Remotion text_card, NOT AI-generated images with text
|
||||
|
||||
@@ -12,7 +12,7 @@
|
||||
| `flux_image` | FLUX 2 Pro via fal.ai | ~$0.03-0.05 | ~5-10s | Photorealism, general purpose, workhorse |
|
||||
| `grok_image` | Grok Imagine Image (xAI) | $0.02/output + $0.002/input edit image | ~5-15s | Image edits, style transfer, multi-image compositing |
|
||||
| `openai_image` | GPT Image 1 (OpenAI) | ~$0.01-0.17 | ~5-15s | Complex instructions, text in images, multi-element |
|
||||
| `recraft_image` | Recraft V4 via fal.ai | ~$0.04-0.25 | ~5-10s | Logos, SVG vectors, brand assets, text rendering |
|
||||
| `recraft_image` | Recraft V4 via fal.ai | ~$0.04-0.25 | ~5-10s | Logos, SVG vectors, brand assets, text rendering (see caveat below) |
|
||||
| `local_diffusion` | Stable Diffusion (local) | Free | ~30s+ | Offline, privacy, free |
|
||||
| `image_gen` | Multi (legacy, deprecated) | Varies | Varies | **Deprecated** — use `image_selector` or per-provider tools |
|
||||
|
||||
@@ -46,6 +46,12 @@
|
||||
| **Budget/free project** | `pexels_image` or `pixabay_image` | Free, immediate | `local_diffusion` |
|
||||
| **Offline/air-gapped** | `local_diffusion` | No network needed | — |
|
||||
|
||||
## Provider-Specific Caveats
|
||||
|
||||
### Recraft V4 via fal.ai
|
||||
- **`style` parameter causes 422 errors** (as of 2026-04). The `style` enum values (`digital_illustration`, `realistic_image`, etc.) are rejected by fal.ai's Recraft V4 endpoint. **Workaround:** encode style direction in the prompt text instead (e.g. "digital illustration of a tooth cross-section" rather than `style="digital_illustration"`). The `image_size` and `colors` parameters work fine.
|
||||
- **Text rendering is unreliable for exact business names.** Recraft (like all AI image models) may hallucinate wrong text. For any scene where text must be verbatim (CTA screens, business names, phone numbers), use Remotion `text_card` instead of generating an image with text.
|
||||
|
||||
## Cost-Quality Tradeoff
|
||||
|
||||
```
|
||||
|
||||
@@ -210,6 +210,7 @@ the AI model's training data — it may be wrong or outdated.
|
||||
- **Ignoring narration timing**: If TTS produces 12s of audio for a 10s section, the edit phase will struggle. Check durations.
|
||||
- **Missing pronunciation guide**: "PostgreSQL" or "Kubernetes" will be mispronounced without explicit guidance.
|
||||
- **One retry then give up**: If an image doesn't match, refine the prompt specifically — don't just retry the same prompt.
|
||||
- **AI-generating images with exact text (CTA, business names, contact info)**: AI image models frequently hallucinate wrong text — wrong business name, wrong phone number, misspelled words. **Never use AI image generation for scenes where text must be verbatim.** Use Remotion `text_card` type instead. This applies to: CTA screens, title cards with business names, contact info overlays, legal disclaimers. If a scene's `type` is `text_card` in the scene plan, do NOT generate an image for it — skip it and let the compose stage render it natively in Remotion.
|
||||
|
||||
|
||||
## When You Do Not Know How
|
||||
|
||||
@@ -22,19 +22,24 @@ This is the last technical stage before the video exists as a playable file. Eve
|
||||
|
||||
Based on the edit decisions, pick the rendering approach:
|
||||
|
||||
**FFmpeg pipeline** (simpler videos):
|
||||
**Remotion render** (DEFAULT — use this unless explicitly overridden):
|
||||
- Animated text cards, stat cards, chart scenes
|
||||
- Complex transitions (morph, zoom, ken-burns)
|
||||
- Programmatic motion graphics
|
||||
- Audio embedding (narration + music with fade/volume)
|
||||
- Word-level captions via CaptionOverlay component
|
||||
- Best for: ALL explainer videos, both image-based and animation-heavy
|
||||
|
||||
**FFmpeg pipeline** (FALLBACK — only when Remotion is unavailable):
|
||||
- Static images with Ken Burns
|
||||
- Audio layering
|
||||
- Subtitle burn-in
|
||||
- Best for: diagram-heavy, image-based explainers
|
||||
- SRT subtitle burn-in
|
||||
- Best for: environments without Node.js/Remotion installed
|
||||
|
||||
**Remotion render** (motion-heavy videos):
|
||||
- Animated text cards, stat cards
|
||||
- Complex transitions (morph, zoom)
|
||||
- Programmatic motion graphics
|
||||
- Best for: flat-motion-graphics playbook, animation-heavy plans
|
||||
|
||||
You can combine both: Remotion for animated segments, FFmpeg for final assembly.
|
||||
**IMPORTANT: When using Remotion, ALL of these go through Remotion — not FFmpeg:**
|
||||
- Audio (narration + music) → Remotion `audio` prop, NOT external audio_mixer
|
||||
- Subtitles → Remotion `captions` prop (word-level), NOT SRT burn via FFmpeg
|
||||
- Text overlays (CTA, titles) → Remotion `text_card` cut type, NOT AI-generated images
|
||||
|
||||
### Step 2: Audio Acquisition (Narration, Music, Subtitles)
|
||||
|
||||
@@ -170,19 +175,34 @@ for the proven formula — especially the all-dark-background rule for visual co
|
||||
|
||||
### Step 5: Audio Post-Processing
|
||||
|
||||
**Remotion path (DEFAULT):** Skip external audio mixing entirely. Remotion handles all audio
|
||||
natively via `<Audio>` components. Pass audio sources in the composition props:
|
||||
```json
|
||||
{
|
||||
"audio": {
|
||||
"narration": { "src": "project/narration.mp3", "volume": 1.0 },
|
||||
"music": { "src": "project/music.mp3", "volume": 0.12, "fadeInSeconds": 1.5, "fadeOutSeconds": 2.5 }
|
||||
}
|
||||
}
|
||||
```
|
||||
Remotion renders audio and video in a single pass — no external muxing needed.
|
||||
Do NOT use `audio_mixer` for ducking/mixing when rendering via Remotion.
|
||||
|
||||
**FFmpeg fallback (ONLY when Remotion is unavailable):**
|
||||
Call the `audio_mixer` tool to:
|
||||
1. Layer narration segments in order
|
||||
2. Mix background music at playbook volume
|
||||
3. Apply ducking (music dips during narration)
|
||||
4. Normalize overall audio levels
|
||||
5. Output the final mixed audio track
|
||||
|
||||
The video_compose tool will mux this with the video.
|
||||
|
||||
### Step 5b: Generate Subtitles (Mandatory)
|
||||
|
||||
Subtitles are mandatory for all explainer content. Generate them from the narration audio — do NOT skip this step.
|
||||
|
||||
**Remotion path (DEFAULT — when using Remotion render):**
|
||||
|
||||
1. **Transcribe** the full narration using the `transcriber` tool (whisperx):
|
||||
```python
|
||||
from tools.analysis.transcriber import Transcriber
|
||||
@@ -195,39 +215,50 @@ Subtitles are mandatory for all explainer content. Generate them from the narrat
|
||||
# result.data contains segments with word-level timestamps
|
||||
```
|
||||
|
||||
2. **Generate SRT** from the transcription using `subtitle_gen`:
|
||||
2. **Convert to Remotion WordCaption format** (NOT SRT):
|
||||
```python
|
||||
captions = []
|
||||
for segment in result.data['segments']:
|
||||
for word_info in segment.get('words', []):
|
||||
captions.append({
|
||||
'word': word_info['word'],
|
||||
'startMs': int(word_info['start'] * 1000),
|
||||
'endMs': int(word_info['end'] * 1000),
|
||||
})
|
||||
```
|
||||
|
||||
3. **Add captions to composition props** — they go in the `captions` array alongside `cuts` and `audio`:
|
||||
```json
|
||||
{
|
||||
"cuts": [...],
|
||||
"audio": {...},
|
||||
"captions": [
|
||||
{ "word": "Root", "startMs": 120, "endMs": 340 },
|
||||
{ "word": "canals", "startMs": 340, "endMs": 680 }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Remotion's CaptionOverlay renders these as word-by-word highlighted captions with the theme's
|
||||
`captionHighlightColor` and `captionBackgroundColor`. This is superior to FFmpeg SRT burn because
|
||||
it produces animated word-level highlighting synchronized to narration.
|
||||
|
||||
**FFmpeg fallback (ONLY when Remotion is unavailable):**
|
||||
|
||||
If Remotion is not available, fall back to SRT generation + FFmpeg burn:
|
||||
```python
|
||||
from tools.subtitle.subtitle_gen import SubtitleGen
|
||||
result = SubtitleGen().execute({
|
||||
SubtitleGen().execute({
|
||||
'segments': transcription_data['segments'],
|
||||
'format': 'srt',
|
||||
'output_path': 'projects/<project>/assets/subtitles.srt',
|
||||
'max_words_per_cue': 8,
|
||||
'max_chars_per_line': 42
|
||||
})
|
||||
# Then burn with video_compose operation='burn_subtitles'
|
||||
```
|
||||
|
||||
3. **Burn subtitles** into the video using `video_compose`:
|
||||
```python
|
||||
from tools.video.video_compose import VideoCompose
|
||||
result = VideoCompose().execute({
|
||||
'operation': 'burn_subtitles',
|
||||
'input_path': 'projects/<project>/renders/output.mp4',
|
||||
'output_path': 'projects/<project>/renders/final.mp4',
|
||||
'subtitle_path': 'projects/<project>/assets/subtitles.srt',
|
||||
'subtitle_style': {
|
||||
'font': '<from playbook typography.headings.font or Arial>',
|
||||
'font_size': 22,
|
||||
'primary_color': '&HFFFFFF',
|
||||
'outline_color': '&H000000',
|
||||
'outline_width': 2,
|
||||
'margin_v': 50,
|
||||
'alignment': 2
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
**The final deliverable is the subtitled version**, not the pre-subtitle render.
|
||||
**The final deliverable MUST have subtitles** — either via Remotion captions or FFmpeg burn.
|
||||
|
||||
### Step 5c: Pre-Render Validation (Mandatory)
|
||||
|
||||
@@ -250,14 +281,30 @@ Common catches:
|
||||
|
||||
**Do not skip this step.** If validation fails, fix the issue and re-validate before rendering.
|
||||
|
||||
### Step 6: Post-Render Self-Review (Mandatory)
|
||||
### Step 6: Post-Render Self-Review (Mandatory — ALL steps required)
|
||||
|
||||
After rendering, the agent **must review its own output** before presenting to the user. This catches issues the validator can't see (visual quality, audio sync, subtitle readability).
|
||||
|
||||
**6a. Extract review frames:**
|
||||
**CRITICAL: You MUST complete ALL of steps 6a through 6e. Do NOT skip any step.
|
||||
The most common agent failure is doing 6a (frames) and 6c (visual) while skipping
|
||||
6b (audio transcription) — which misses catastrophic issues like missing audio entirely.**
|
||||
|
||||
**6a. Probe rendered file (FIRST — gate for all other checks):**
|
||||
```bash
|
||||
ffprobe -v quiet -print_format json -show_format -show_streams rendered_video.mp4
|
||||
```
|
||||
Verify:
|
||||
- Video stream exists (codec_type: "video") with correct resolution
|
||||
- **Audio stream exists (codec_type: "audio")** — if NO audio stream, STOP and fix immediately
|
||||
- Duration is within ±5% of target
|
||||
- File size is reasonable (not 0 bytes)
|
||||
|
||||
**If audio stream is missing: the render did not embed audio. Do NOT proceed to present
|
||||
the video to the user. Fix the audio configuration and re-render.**
|
||||
|
||||
**6b. Extract review frames:**
|
||||
```python
|
||||
from tools.analysis.frame_sampler import FrameSampler
|
||||
# Extract one frame per scene at the midpoint
|
||||
midpoints = [(cut['in_seconds'] + cut['out_seconds']) / 2 for cut in cuts]
|
||||
FrameSampler().execute({
|
||||
'input_path': 'path/to/rendered_video.mp4',
|
||||
@@ -268,37 +315,41 @@ FrameSampler().execute({
|
||||
})
|
||||
```
|
||||
|
||||
**6b. Transcribe rendered audio:**
|
||||
**6c. Transcribe rendered audio (MANDATORY — do NOT skip):**
|
||||
```python
|
||||
from tools.analysis.transcriber import Transcriber
|
||||
Transcriber().execute({
|
||||
result = Transcriber().execute({
|
||||
'input_path': 'path/to/rendered_video.mp4',
|
||||
'model_size': 'base',
|
||||
'language': 'en',
|
||||
'output_dir': 'path/to/review-frames',
|
||||
})
|
||||
# Verify all narration words are present and not cut off
|
||||
# If result returns 0 words: audio is silent/missing — STOP and fix
|
||||
# If word count < 80% of script word count: audio is cut off — investigate
|
||||
```
|
||||
|
||||
**6c. Visual inspection — review each frame:**
|
||||
**6d. Visual inspection — review each frame:**
|
||||
- Does the background color/gradient match intent? (watch for white backgrounds on dark-themed videos)
|
||||
- Are images rendering correctly? (not blank, not stretched)
|
||||
- Are subtitles visible and properly spaced?
|
||||
- Are subtitles/captions visible and properly spaced?
|
||||
- Are overlays (section titles, stat reveals) positioned correctly?
|
||||
- Is the opening scene visually strong? (important for social media thumbnails)
|
||||
- Does the CTA/closing screen show correct text? (AI-generated text in images frequently hallucinates — use Remotion text_card for any text that must be exact)
|
||||
|
||||
**6d. Audio inspection — check transcript:**
|
||||
**6e. Audio inspection — check transcript against script:**
|
||||
- Is the full narration captured? (compare last transcribed word to last scripted word)
|
||||
- Any words cut off at the end? (narration exceeding video duration)
|
||||
- Timing alignment — do narration segments roughly match their intended scenes?
|
||||
- Is background music audible? (transcriber may not capture music, but ffprobe confirms audio stream)
|
||||
|
||||
**6e. Compile and present review to user:**
|
||||
**6f. Compile and present review to user:**
|
||||
|
||||
> **Post-render review for "[Video Title]":**
|
||||
>
|
||||
> **Audio:** [Complete/Cut off at Xs] — all N words captured / last sentence missing
|
||||
> **File:** [duration]s, [resolution], [file size] — audio stream: [present/MISSING]
|
||||
> **Audio:** [Complete/Cut off at Xs] — [N]/[M] words transcribed from rendered output
|
||||
> **Visuals:** [N scenes inspected] — [issues or "all scenes rendering correctly"]
|
||||
> **Subtitles:** [Present/Missing] — [spacing ok / words running together]
|
||||
> **Captions:** [Remotion CaptionOverlay / FFmpeg SRT / MISSING] — [word-level highlight working / issues]
|
||||
> **Issues found:** [list any issues with severity]
|
||||
>
|
||||
> **Recommendations:** [what to fix, if anything]
|
||||
|
||||
@@ -223,3 +223,4 @@ Call `handle_explainer_scene_plan(state, {"scene_plan": scene_plan_json})` to va
|
||||
- **Vague required_assets**: "An image about databases" is useless for prompt engineering. "Isometric illustration of a vector database with embedding vectors floating in 3D space, using the playbook's blue-green palette" is actionable.
|
||||
- **Preset thinking**: A scene plan that says "make it flat-motion-graphics" is not enough. The planner must specify what makes THIS video's motion graphics feel distinct.
|
||||
- **Static scenes for dynamic concepts**: If the narrator describes a process or transformation, the visual should move. Use animation or progressive reveal, not a static image.
|
||||
- **Using `generated` type for CTA/closing screens with exact text**: AI image models hallucinate text — wrong business names, misspelled words, wrong phone numbers. Any scene with verbatim text (CTA, business info, contact details, legal) MUST be `type: "text_card"` so Remotion renders the text exactly. Never plan a `generated` image for a scene where text accuracy matters.
|
||||
|
||||
@@ -144,6 +144,14 @@ class RecraftImage(BaseTool):
|
||||
if inputs.get("image_size"):
|
||||
payload["image_size"] = inputs["image_size"]
|
||||
if inputs.get("style"):
|
||||
# NOTE: As of 2026-04, fal.ai's Recraft V4 endpoint rejects the
|
||||
# `style` parameter with a 422 Unprocessable Entity error. The
|
||||
# style enum values (digital_illustration, realistic_image, etc.)
|
||||
# are NOT accepted by the /fal-ai/recraft/v4/text-to-image route.
|
||||
# Workaround: encode the style direction in the prompt text instead
|
||||
# (e.g. "digital illustration of..." rather than style="digital_illustration").
|
||||
# We still pass the parameter through in case fal.ai re-enables it,
|
||||
# but callers should be aware this may fail.
|
||||
payload["style"] = inputs["style"]
|
||||
if inputs.get("colors"):
|
||||
payload["colors"] = inputs["colors"]
|
||||
|
||||
Reference in New Issue
Block a user