Files
OpenMontage/docs/comfyui-adapter-plan.md
T
martimramos 6ec2bbb090 comfyui: add native ComfyUI provider for image, video, and music generation
Adds three new BaseTool providers that delegate GPU work to a running
ComfyUI server via its REST API.  This avoids the need to install
PyTorch/diffusers directly, which is critical on hardware where the
ecosystem hasn't caught up (e.g. NVIDIA Blackwell / DGX Spark, aarch64
+ CUDA 13.0).

New files:
- tools/_comfyui/client.py — shared REST client (submit/poll/download)
- tools/_comfyui/workflows/ — 4 bundled workflow templates
- tools/graphics/comfyui_image.py — FLUX 2 Dev NVFP4 text-to-image
- tools/video/comfyui_video.py — WAN 2.2 14B t2v + i2v (4-step LightX2V)
- tools/audio/comfyui_music.py — ACE-Step 3.5B music generation
- tests/contracts/test_comfyui_tools.py — 41 contract tests
- docs/comfyui-adapter-plan.md — design document

Zero changes to existing tools, selectors, registry, or pipelines.
Tools are auto-discovered and selectors pick them up via capability match.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-06-29 19:08:49 +01:00

13 KiB

ComfyUI Provider Adapter for OpenMontage

RFC: Native ComfyUI backend for image, video, and music generation


Motivation

OpenMontage's local GPU tools (wan_video, hunyuan_video, cogvideo_video, local_diffusion) use HuggingFace diffusers directly. This works on x86 + consumer GPUs but breaks on newer hardware where the PyTorch ecosystem hasn't caught up:

Issue Detail
NVIDIA Blackwell (sm_121) No stable PyTorch wheels for aarch64 + CUDA 13.0. Requires NGC containers or nightly builds.
Flash Attention Does not support sm_121. Must be replaced with SageAttention v3 or native SDPA.
Unified Memory (GB10/DGX Spark) nvidia-smi cannot report VRAM. Diffusers' memory estimation breaks.
Model format mismatch Diffusers expects HF repos. Production deployments use .safetensors checkpoints with quantized variants (NVFP4, FP8) that diffusers doesn't natively load.

ComfyUI already solves all of these. NVIDIA ships official ComfyUI containers for DGX Spark. The community has optimized workflows for Blackwell (SageAttention, NVFP4 quantization, LightX2V 4-step LoRAs). Models like WAN 2.2, FLUX 2, and ACE-Step run reliably through ComfyUI on hardware where diffusers cannot.

A ComfyUI adapter gives OpenMontage access to any model ComfyUI supports, on any hardware ComfyUI runs on, without shipping or maintaining PyTorch builds.


Design

Architecture

OpenMontage Agent
    |
    v
video_selector / image_selector / music_selector
    |                                        
    v                                        
comfyui_video    comfyui_image    comfyui_music    (new tools)
    |                |                |
    v                v                v
ComfyUI REST API  (POST /prompt, GET /history, GET /view)
    |
    v
GPU (any hardware ComfyUI supports)

Integration model

Three new BaseTool subclasses plus one shared client library:

tools/
  _comfyui/
    __init__.py
    client.py              # Shared ComfyUI REST client
    workflows/             # Bundled workflow templates
      flux2-txt2img.json
      wan22-t2v-4step.json
      wan22-i2v-4step.json
      ace-step-music.json
  graphics/
    comfyui_image.py       # capability="image_generation", provider="comfyui"
  video/
    comfyui_video.py       # capability="video_generation", provider="comfyui"
  audio/
    comfyui_music.py       # capability="music_generation", provider="comfyui"

Zero changes to selectors or registry

The tools declare capability and provider as class attributes. tool_registry.discover() picks them up automatically via pkgutil.walk_packages. video_selector, image_selector, and music_selector find them via registry.get_by_capability() -- no hardcoded references needed.


Shared Client: tools/_comfyui/client.py

Encapsulates the ComfyUI REST API pattern proven in production (used by the Bard project's Airflow DAGs for thousands of generations):

class ComfyUIClient:
    """Thin client for the ComfyUI REST API."""

    def __init__(self, server_url: str | None = None):
        self.server_url = server_url or os.environ.get(
            "COMFYUI_SERVER_URL", "http://localhost:8188"
        )

    def is_available(self) -> bool:
        """Health check -- can we reach the server?"""

    def submit(self, workflow: dict) -> str:
        """POST /prompt. Returns prompt_id. Raises on node_errors."""

    def poll(self, prompt_id: str, timeout: int = 600, interval: int = 5) -> dict:
        """GET /history/{prompt_id} until complete. Returns outputs dict."""

    def download(self, filename: str, subfolder: str, dest: Path) -> Path:
        """GET /view?filename=...&type=output. Writes bytes to dest."""

    def upload_image(self, local_path: Path, name: str) -> str:
        """POST /upload/image. Returns server-side filename for LoadImage nodes."""

    def generate(self, workflow: dict, output_node: str, dest: Path,
                 timeout: int = 600) -> Path:
        """Full cycle: submit -> poll -> download. Returns artifact path."""

Why a shared client? The submit/poll/download cycle is identical across image, video, and music generation. The only differences are: which workflow template, which nodes to customize, and which output node to read from.


Tool Specifications

comfyui_image -- Image Generation

Field Value
capability image_generation
provider comfyui
runtime LOCAL_GPU
tier GENERATE
stability EXPERIMENTAL
capabilities text_to_image, image_to_image
dependencies (runtime: ComfyUI server reachable)
fallback_tools flux_image, local_diffusion, openai_image
cost $0.00 (local compute)

Bundled workflow: flux2-txt2img.json

Loads FLUX 2 Dev (NVFP4) with Mistral text encoder. Templated nodes:

Node Class Templated field
4 CLIPTextEncode text (prompt)
6 EmptyFlux2LatentImage width, height
7 RandomNoise noise_seed
10 Flux2Scheduler steps
13 SaveImage filename_prefix

Input schema:

prompt:        string    # required
width:         integer   # default 1024
height:        integer   # default 1024
steps:         integer   # default 20
seed:          integer   # optional (random if omitted)
guidance:      number    # default 3.5
output_path:   string    # where to save the image
workflow_json: string    # optional override (full custom workflow)

get_status(): Pings ComfyUI server. Returns AVAILABLE if reachable, UNAVAILABLE otherwise.

execute() flow:

  1. Deep-copy workflow template
  2. Inject prompt, seed, dimensions, steps into templated nodes
  3. client.generate(workflow, output_node="13", dest=output_path)
  4. Return ToolResult with artifact path, seed, model info

comfyui_video -- Video Generation

Field Value
capability video_generation
provider comfyui
runtime LOCAL_GPU
tier GENERATE
stability EXPERIMENTAL
capabilities text_to_video, image_to_video
dependencies (runtime: ComfyUI server reachable)
fallback_tools wan_video, hunyuan_video, ltx_video_local
cost $0.00 (local compute)

Bundled workflows:

  1. wan22-i2v-4step.json -- Image-to-video (WAN 2.2 14B, fp8, 4-step LightX2V LoRA)
  2. wan22-t2v-4step.json -- Text-to-video (WAN 2.2 14B, fp8, 4-step LightX2V LoRA)

I2V workflow -- templated nodes:

Node Class Templated field
93 CLIPTextEncode text (positive prompt)
97 LoadImage image (server filename from upload)
98 WanImageToVideo width, height, length
86 KSamplerAdvanced noise_seed
108 SaveVideo filename_prefix

Input schema:

prompt:               string    # required
operation:            string    # "text_to_video" | "image_to_video" (default: t2v)
reference_image_path: string    # local path (for i2v)
reference_image_url:  string    # URL (for i2v, downloaded first)
width:                integer   # default 640
height:               integer   # default 640
num_frames:           integer   # default 81 (5s at 16fps)
seed:                 integer   # optional
output_path:          string    # where to save the video
workflow_json:        string    # optional override

execute() flow (i2v):

  1. Upload reference image via client.upload_image()
  2. Deep-copy i2v workflow template
  3. Inject prompt, uploaded image name, seed, dimensions
  4. client.generate(workflow, output_node="108", dest=output_path, timeout=900)
  5. Return ToolResult

execute() flow (t2v):

  1. Deep-copy t2v workflow template
  2. Inject prompt, seed, dimensions
  3. client.generate(workflow, output_node="108", dest=output_path, timeout=900)
  4. Return ToolResult

comfyui_music -- Music Generation

Field Value
capability music_generation
provider comfyui
runtime LOCAL_GPU
tier GENERATE
stability EXPERIMENTAL
capabilities text_to_music
dependencies (runtime: ComfyUI server reachable + ACE-Step model)
fallback_tools suno_music, elevenlabs_music
cost $0.00 (local compute)

Bundled workflow: ace-step-music.json

Uses ACE-Step v1 3.5B for text-to-music generation. Workflow to be authored based on the ComfyUI ACE-Step custom node.

Input schema:

prompt:      string    # required (music description)
duration:    number    # seconds (default 30)
seed:        integer   # optional
output_path: string    # where to save the audio

Workflow Override Mechanism

Every tool accepts an optional workflow_json input. When provided, it replaces the bundled template entirely. This enables:

  • Using newer model checkpoints without code changes
  • Custom sampling strategies (different schedulers, step counts, LoRAs)
  • Community workflows dropped in as-is
  • A/B testing different generation approaches

The agent can also read workflow files from tools/_comfyui/workflows/ and modify them programmatically before passing to execute().


Configuration

Environment variables:

# .env
COMFYUI_SERVER_URL=http://localhost:8188    # ComfyUI API endpoint
COMFYUI_POLL_INTERVAL=5                     # seconds between status checks
COMFYUI_POLL_TIMEOUT=600                    # max wait for image gen
COMFYUI_VIDEO_TIMEOUT=900                   # max wait for video gen

For Docker Compose setups (ComfyUI in a container):

COMFYUI_SERVER_URL=http://host.docker.internal:8188
# or
COMFYUI_SERVER_URL=http://comfyui:8188      # if on same docker network

Provider Selection Behavior

When the adapter is available, selectors will rank it alongside other providers using OpenMontage's 7-dimension scoring:

Dimension ComfyUI score Rationale
Task fit High Supports t2i, i2v, t2v, music
Quality High Latest models (FLUX 2, WAN 2.2 14B)
Control Highest Full workflow customization
Reliability High Proven in production
Cost $0 Local compute
Latency Medium GPU-bound, no network round-trip
Continuity High Deterministic with seeds

When ComfyUI is unavailable (server down), the selector falls through to fallback_tools automatically -- API providers like FLUX via fal.ai or HeyGen take over transparently.


What This Unlocks

Immediate (with existing models)

  • FLUX 2 Dev NVFP4 image generation -- Blackwell-optimized, ~60s per image
  • WAN 2.2 14B i2v with 4-step acceleration -- ~3.5 min per 5s clip
  • WAN 2.2 14B t2v (models downloaded, workflow needed)
  • ACE-Step 3.5B local music generation (model downloaded, workflow needed)

Future (add models to ComfyUI, no code changes to OpenMontage)

  • Newer checkpoints (WAN 3.x, FLUX 3, etc.) -- just update workflow JSON
  • ControlNet, IP-Adapter, AnimateDiff -- supported via ComfyUI custom nodes
  • Upscaling, inpainting, outpainting -- ComfyUI nodes exist
  • Any model the ComfyUI ecosystem supports

Hardware portability

The same adapter works on:

  • NVIDIA DGX Spark (GB10, aarch64, CUDA 13.0)
  • Consumer GPUs (RTX 3090/4090, x86)
  • Cloud instances (A100, H100)
  • Multi-GPU setups (ComfyUI handles device placement)

No PyTorch version pinning, no architecture-specific wheels, no CUDA compatibility matrices. ComfyUI is the abstraction layer.


Implementation Scope

Component Files Estimated size
Shared client tools/_comfyui/client.py ~120 lines
Image tool tools/graphics/comfyui_image.py ~130 lines
Video tool tools/video/comfyui_video.py ~160 lines
Music tool tools/audio/comfyui_music.py ~100 lines
Workflow templates tools/_comfyui/workflows/*.json 4 files
T2V workflow tools/_comfyui/workflows/wan22-t2v-4step.json 1 file (to author)
Music workflow tools/_comfyui/workflows/ace-step-music.json 1 file (to author)
Tests tests/contracts/test_comfyui_*.py ~80 lines
Docs skills/creative/comfyui-workflows.md Agent skill file

Total: ~600 lines of Python + 4-6 workflow JSONs.

No changes to: base_tool.py, tool_registry.py, any selector, any existing tool, any pipeline definition, or any schema.


Open Questions

  1. Workflow versioning: Should workflow JSONs live in the repo or be user-provided via a config directory? Bundling gives reproducibility; external gives flexibility.

  2. Model discovery: ComfyUI has a /object_info endpoint that lists available nodes and models. Should get_status() also report which models are loaded, so the selector can make informed routing decisions?

  3. Async generation: ComfyUI supports websocket connections for real-time progress. Worth implementing for long video generations, or is polling sufficient?

  4. Multi-server: Should the adapter support multiple ComfyUI instances (e.g., one for images, one for video) via per-capability URLs?