Anyone who has shipped something built on hosted generative models knows the actual problem, and it isn’t model quality. It’s that you integrated five providers, each with its own SDK, key, auth flow, rate limits and failure behaviour — and when one of them degrades at 2am you find out from a user.
Comfy Router, live since 23 September 2026, is ComfyUI’s answer: one API for frontier image, video, 3D and audio models.
What it does
Same model string. Same arguments. No new SDK, no new key, no redeploy.
Day-one coverage:
- Seedance 2.5
- MiniMax H3
- Nano Banana Pro
- GPT Image 2
- Kling
- Black Forest Labs
with more stated as coming. You integrate once, then choose the provider per job — for availability, or for price.
The design decision worth noticing
Most routers sell automatic failover as the headline feature. Comfy Router deliberately doesn’t. From their own description: explicit routing — you name the provider, they call that provider. “It’s down? The request fails there.”
That sounds like a downgrade and is the opposite. Silent failover is a debugging nightmare in generative media specifically, because different models produce different output. If your pipeline quietly reroutes from one image model to another during an outage, you don’t get an error — you get a batch of assets that look wrong, discovered later, with no record of why. A router that fails loudly is a router you can reason about.
The corollary is the genuinely useful bit: every job returns the provider that ran it, so you can log, bill and debug per request. Cost attribution across five providers is otherwise a spreadsheet exercise.
The async model fits generative media
The API is asynchronous by design: submit() returns a request ID immediately, the queue retries errors until a slot opens, and subscribe() submits and polls to completion.
This matters because video generation is slow and capacity-constrained in a way text generation isn’t. A synchronous call that blocks for four minutes and then times out is the wrong abstraction. A request ID you can poll, with queue-level retry when the provider is simply full, matches how these models actually behave under load.
Who this is and isn’t for
It is for anyone building a product, tool or pipeline on hosted generative media — where provider outages, price differences and cost attribution are operational problems you’re already living with.
It isn’t for people running models locally. If your workflow is Qwen-Image GGUF on your own GPU in the ComfyUI you already have, a router to hosted frontier models is orthogonal to what you’re doing.
The strategic read: ComfyUI has been the open-source local-generation environment, and this is a deliberate move into being infrastructure for the hosted side too. Whether that’s welcome depends on whether you think the project’s centre of gravity should stay local. The engineering choices here — explicit routing, provider attribution, async queueing — are the ones you’d make if you actually ran production workloads, which is a good sign about who designed it.