Current diffusion fashions allow high-quality video technology, however undergo from gradual runtimes. The big transformer-based backbones utilized in these fashions are bottlenecked by spatiotemporal consideration. On this paper, we determine {that a} important fraction of token-to-token connections constantly yield negligible scores throughout numerous inputs, and their patterns usually repeat throughout queries. Thus, the eye computation in these instances could be skipped with little to no impact on the end result. This statement continues to carry for connections amongst native token blocks. Motivated by this, we introduce CalibAtt, a training-free technique that accelerates video technology through calibrated sparse consideration. CalibAtt performs an offline calibration go that identifies block-level sparsity and repetition patterns which are secure throughout inputs, and compiles these patterns into optimized consideration operations for every layer, head, and diffusion timestep. At inference time, we compute the chosen input-dependent connections densely, and skip the unselected ones in a hardware-efficient method. In depth experiments on Wan 2.1 14B, Mochi 1, and few-step distilled fashions at numerous resolutions present that CalibAtt achieves as much as 1.58× end-to-end speedup, outperforming present training-free strategies whereas sustaining video technology high quality and text-video alignment.
- †Tel Aviv College
- ** Work performed whereas at Apple







