What a 5-second AI video clip costs on a rented GPU
Hosted video generators charge somewhere between $0.25 and $0.40 for a 5-second 720p clip. Open-weight models can run on a GPU you rent by the hour. On 29 September 2026 we spent $0.46 finding out what that costs in practice: the clip prices are low, but the first clip of a session is not cheap in time, and the cheap card has a hard ceiling.
The measurements
One community-cloud RTX 4090 (24 GB) at $0.34 an hour, 5-second clips at about 1280 by 720. One run per model. GPU time is the seconds per clip multiplied by the hourly price; for example 483 s × $0.34 / 3600 = $0.0456.
| Model | Seconds per clip | GPU time per clip | Including cold start | Output |
|---|---|---|---|---|
| Wan 2.2 A14B 4-step, text-to-video | 483 | $0.046 | $0.063 | 16 fps, 81 frames |
| Wan 2.2 A14B 4-step, image-to-video | 501 | $0.047 | $0.064 | 16 fps, 81 frames |
| Wan 2.2 TI2V-5B | 436 | $0.041 | $0.065 | 24 fps, 121 frames |
| LTX-2.5 distilled, text-to-video | 275 | $0.026 | $0.034 | 24 fps, stereo audio |
| LTX-2.5 distilled, image-to-video | 277 | $0.026 | $0.035 | 24 fps, stereo audio |
Image-to-video was within 4% of text-to-video for both families. There were no failures, out-of-memory kills or black frames. We did not score quality: the clips were not rated, so this page says nothing about which one looks better.
Against the hosted bar of $0.25–0.40 per clip (list prices we collected in desk research, not measured by us), GPU time alone is roughly 5 to 15 times cheaper. That comparison ignores your own time, idle hours and reruns.
The cold start is the real session cost
- Container image pull: 38 s. Setup: 191 s.
- Model weights: 294 s for Wan 14B (62 GB), 177 s for LTX (39 GB), 24 s for the Wan 5B model.
- Pod creation to first finished clip: about 7.5 minutes (LTX) to 10 minutes (Wan 14B).
Spread over a session, that moves the cost per clip from 4.6 to 6.3 cents for Wan 14B and from 2.6 to 3.4 cents for LTX when the cold start is spread over the clips of that session. Over 100 clips it nearly disappears (4.6 and 2.6 cents). Weights need a cache, either baked into the image or on a volume, otherwise every cold pod pays minutes of downloading before it earns anything.
Why the 4090 is a floor, not a target
- The card is full. VRAM peaked at 23.5 GiB for every model.
- It was slower than our desk research predicted. Wan 14B took about 108 s per step at 720p; the estimate we had read was about 150 s per clip.
- LTX mostly shuffles memory. It generated for only 78 of its 275 seconds; the rest was swapping components in and out of 24 GB. We expect a 32–48 GB card to be about twice as fast for LTX, but that is an expectation, not a run.
- Container memory is the real limit. The pod’s 50 GB cgroup, not the host’s 251 GB, is what counts: the ComfyUI workflow tool plus a parallel download was OOM-killed. Never download while generating.
- Community capacity comes and goes. Stock of the 4090 appeared and disappeared minute by minute.
What we did not measure
MiniMax-H3 was not measured. It needs an H100-class card, and we found no community H100 at or under $2.20 an hour in 45 attempts over about 18 minutes, so we did not rent one. Also not measured: 10-second clips, repeat runs, several prompts per model, quality scores, 32–48 GB cards, and a persistent volume against a baked image. Treat the numbers as one careful sample, not a benchmark.
Licences (public information, not legal advice)
This is from the licence texts and vendor pages as we read them in September 2026, not from our testing, so verify before you rely on it.
- Wan 2.2: Apache 2.0, commercial use allowed.
- LTX-2.x: the Lightricks community licence is free below $10 million in annual revenue, with a paid agreement at or above it. Revenue-generating use counts as commercial.
- MiniMax-H3: its own community licence, with a commercial-application form for the US, EU, UK and South Korea.
What this means for how we are building it
The shape we are working toward is a GPU pod that belongs to the customer, rather than a shared pool, so the cold start and the idle hours are visible instead of hidden in a per-clip price. Given the capacity result above, that pod would run on a provider tier with reliable stock rather than the community market. This is a design direction, not a launched product. For the rest of the GPU picture, see why GPU containers break.
Need a server for the agent that drives it?
The agent that orchestrates a video pipeline runs fine on a small CPU server.
Related
- Why your GPU container breaks — CUDA, drivers and pinned tags.
- Run Blender headless on a VPS — the CPU-only side of 3D.
- What an AI agent turn costs — the same per-unit thinking for agent chat.
Measured on
- Community-cloud RTX 4090 24 GB at $0.34/h, 29 September 2026; three pods, all terminated afterwards.
- Total spend $0.46 of a $4.00 budget.
One run per model, 5-second clips at about 720p, no quality scoring. Hosted prices and licence terms are third-party information, not our measurement.