Realtime Video Generation Calculator
First principles estimate — adjust parameters, get concrete numbers
Model
Hardware
Target Video
How the estimate works
The calculator estimates the FLOPs required to generate one second of video from first principles, then divides by the effective throughput of the selected hardware. No black-box magic — every number on the page comes from the formulas below.
Token count
Video is processed as latent tokens, not pixels. A VAE compresses each frame spatially (8×8 in Wan2.1, 16×16 in Wan2.2), and a patch size of 2 further halves each dimension, so a 720p frame becomes 1280/32 × 720/32 = 40 × 22 tokens per latent frame with 16×16 compression. Time is compressed 4× by the causal 3D-VAE: one second at 16 FPS becomes 1 + ⌊(16−1)/4⌋ = 4 latent frames. Total tokens per second of video = latT × latH × latW.
FLOPs per denoising step
Each DiT layer contributes three terms:
- Self-attention:
8·T·d² + 4·T²·d, where T is tokens and d is the hidden dim. The quadraticT²term is what makes long, high-resolution video expensive — and it is the term that sparse attention divides (2×–6×). - Cross-attention to the text prompt: calibrated to ~8.4 TFLOP per layer on a 14B model, scaled linearly with model size.
- FFN:
4·T·d·d_ff— the largest linear term for typical sequence lengths.
Hidden dimension and FFN size scale as √(params/14B) from the Wan2.1 14B reference; layer counts follow known architectures (24 layers at 1.3B up to 80 at 200B). Per-step FLOPs are then multiplied by the number of denoising steps — this is why distillation from 50 steps down to 4 is a ~12× speedup — and doubled when CFG is on, since classifier-free guidance runs both a conditional and an unconditional pass.
Hardware throughput
Peak TFLOPS of the selected GPU at the chosen precision are multiplied by 40% MFU (model FLOPs utilization — the share of peak compute a well-optimized inference stack actually achieves). Multi-GPU setups apply a 0.7 scaling-efficiency factor for communication overhead. Mobile NPUs are converted at roughly TOPS × 0.15 to FP16-equivalent TFLOPS.
AR streaming models
Autoregressive streaming models (e.g. LongLive) generate frame-by-frame with a local attention window, so cost scales ~linearly with pixels rather than quadratically. The estimate is calibrated to the published LongLive 1.3B benchmark — 24.8 FPS at 480p on a single H100 in FP8, i.e. ~32 TFLOP per frame — and scaled by resolution and quantization.
Limitations
These are order-of-magnitude engineering estimates, not benchmarks. Memory-bandwidth-bound regimes (small batch, attention decode), kernel efficiency of specific stacks (FlashAttention, SageAttention, compile), and VAE decode cost are not modeled. Treat results within roughly ±2× as "correct".
FAQ
When will realtime AI video generation be possible?
It already is — for small distilled models at 480p. A 1.3B model at 4 steps runs in realtime on a single H100 today, and AR streaming models like LongLive hit 20+ FPS. The frontier question is when 14B-class quality becomes realtime: on this calculator's numbers, that needs roughly a B200/Rubin-class GPU at 4 steps with sparse attention — i.e. 2025–2026 hardware.
What is MFU and why 40%?
MFU (model FLOPs utilization) is the fraction of a GPU's theoretical peak compute that real workloads sustain. Attention softmax, memory transfers, and kernel launch gaps mean even well-optimized inference rarely exceeds 40–50% of peak. Using peak numbers directly would overstate speed by 2.5×.
Why do denoising steps dominate the cost?
Each step is a full forward pass through the model, and steps multiply total compute linearly. Going from 50 steps (typical for quality) to 4 steps (distilled / few-step models like Turbo or Lightning variants) is a 12.5× speedup with no hardware change — the single largest lever after model size.
Can video generation run on a phone NPU?
Not yet for diffusion at usable quality: a 80-TOPS mobile NPU converts to roughly 12 effective FP16 TFLOPS, which is 1–2 orders of magnitude short of realtime even for a distilled 1.3B model at 480p. AR streaming with INT4 quantization is the more plausible mobile path — hence the INT8/INT4 options above.
How accurate are these numbers?
Within about ±2× of measured inference on optimized stacks. The model is calibrated against public benchmarks (Wan2.1/Wan2.2 architecture parameters, LongLive's published 24.8 FPS on H100) and extrapolated from there. Use it to compare configurations and hardware generations, not to budget a production deployment.