Commit Graph

14 Commits

Author SHA1 Message Date
calesthio a919fde450 Remove fal-first provider guidance 2026-04-08 13:04:32 -07:00
calesthio 7b40edb82c Fix gap #22 and add CinematicRenderer captions + music support
- AGENT_GUIDE.md: add "never read source code" rule — skills are the
  interface, not .py files
- animation.yaml: add Layer 2 skill-first guardrail in assets stage
- video-reference-analyst.md: Step 4b now mandates Layer 2 before Layer 3,
  explicitly forbids reading implementation code
- CinematicRenderer: add TikTok-style CaptionOverlay and separate music
  track support (narration + music as independent audio layers)
- cinematic/types.ts: add CinematicCaptionConfig and music prop types
2026-04-04 14:23:46 -07:00
calesthio 370f2f11b7 Default TTS recommendation to Google Chirp3-HD over ElevenLabs
Chirp3-HD is near-free ($0.003 vs $0.30+ per script), expressive,
24kHz, and available without quota limits. ElevenLabs demoted to
voice-cloning-only recommendation.
2026-04-04 14:08:34 -07:00
calesthio fc6ff2248a Add UAT guardrails to pipeline YAML and fix playbook test
- animation.yaml: enforce audio architecture + provider comparison in
  proposal stage, Layer 3 skill gate + clip duration in assets stage
- video-reference-analyst.md: Step 6 is now a hard redirect forcing
  stage-by-stage pipeline execution
- Fix test_compatible_with_manifest for nested compatible_playbooks dict
2026-04-04 13:29:24 -07:00
calesthio 93de8effbe Add mandatory Layer 3 skill gate before any asset generation
New Step 4b requires the agent to read agent_skills from every tool
before writing generation prompts. Discovered during UAT: agent
generated video clips, images, and TTS without reading provider-specific
prompting guidance, violating AGENT_GUIDE governance.
2026-04-04 13:20:19 -07:00
calesthio 97097e61a1 Present video gen provider options with costs to user, don't auto-pick
Agent must show a provider comparison table (quality, speed, cost per
clip) and recommend one, but let the user choose.
2026-04-04 13:18:15 -07:00
calesthio 1451f4fd69 Add clip duration optimization and Layer 3 skill reminder to proposals
Proposals must now specify clip duration strategy (prefer 10s over 5s
to halve API costs) and list Layer 3 skills that must be read before
generating assets. Both were governance gaps found during UAT.
2026-04-04 13:17:31 -07:00
calesthio 24a85337c3 List all TTS providers in proposal template, not just ElevenLabs
Agent must run tts_selector preflight to check available providers
instead of defaulting to ElevenLabs. Supports ElevenLabs, Google TTS,
OpenAI TTS, and Piper (offline).
2026-04-04 12:53:34 -07:00
calesthio cb1423a61f Add audio architecture decision to video-reference-analyst planning
The agent must lock the audio approach (single narrator vs. character
dialogue vs. both) during Step 3, not defer it to script/compose stage.
Proposals now include voice casting with specific voice IDs.
2026-04-04 12:51:47 -07:00
calesthio 16647d2d36 Enforce Remotion-first composition engine, fix FFmpeg fallback bugs
Remotion is now the default composition engine for ALL final renders
when available — video clips, images, mixed content. FFmpeg is only
used as fallback when Remotion is not installed or for standalone
operations (trim, transcode). Also fixes three FFmpeg fallback bugs:
profile + copy codec conflict, stream order mapping, and segment
seeking for audio-first containers.
2026-04-04 12:39:24 -07:00
calesthio 38b6ec5212 Add mandatory lightweight research step to video-reference-analyst
The reference analyst skill went straight from capability audit to
creative proposals with no research. This caused the agent to propose
concepts based solely on the reference analysis and its own knowledge,
missing content landscape context, technique best practices, and
subject-matter depth.

Added Step 3b between critical questions and creative proposals:
- Content landscape scan (3-5 similar existing videos)
- Style/technique research (AI model strengths, prompting patterns)
- Subject-matter research (facts, tropes, hooks)
- 2-3 minute time budget — lightweight, not full research-director
2026-04-04 11:59:11 -07:00
calesthio 61c704149e Enforce Remotion-first composition in video-reference-analyst skill
The AGENT_GUIDE establishes Remotion as the preferred composition engine
over FFmpeg, but the reference analyst skill was presenting them as peer
options. This caused the agent to default to FFmpeg during proposals.

- Capability audit template now labels Remotion as "preferred" and FFmpeg
  as "fallback only"
- Added explicit composition engine priority note
- Proposal template separates Composition and Motion as distinct lines
2026-04-04 11:43:17 -07:00
calesthio 286c26e33d Add per-scene motion classification to video analyzer
Video analyzer now uses Farneback dense optical flow to classify each
scene as motion_clip, animated_still, or static_image. This lets the
agent correctly identify whether a reference video uses AI-generated
video clips vs still images with pan/zoom — and plan the right pipeline.

Changes:
- video_analyzer.py: new Step 3b with _classify_scene_motion() and
  _read_frame_at() helpers; updated _needs_motion() to use per-scene
  motion data instead of pacing heuristic alone
- video-reference-analyst.md: added Motion line to summary template
  and instructions to read motion_type field before proposing tools
2026-04-04 11:33:33 -07:00
calesthio b0917d2d84 Add reference video input analysis workflow 2026-04-04 10:01:11 -07:00