Unified multimodal fashions that perceive, cause over, and generate interleaved textual content–picture sequences stay structurally fragmented: current approaches both sacrifice visible constancy by means of discrete tokenization, impose structural asymmetry by combining causal textual content era with iterative diffusion-based denoising, or degrade pretrained understanding when adapting vision-language fashions for era. We observe that autoregressive normalizing flows are autoregressive Transformers—sharing the identical causal masks, KV-cache mechanism, and left-to-right construction as LLMs—making them essentially the most pure paradigm for really unified multimodal era that’s steady, single-pass, and purely causal. We current STARFlow2, constructed on the Pretzel structure that vertically interleaves a frozen pretrained VLM stream with a TARFlow stream through residual skip connections, each working below the identical causal masks. This design concurrently preserves pretrained multimodal understanding, allows high-fidelity steady picture era, and achieves structural unification below a single causal mechanism. Mixed with a deep-shallow movement design and a unified FAE latent area, STARFlow2 helps cache-friendly interleaved era the place each textual content and visible outputs instantly enter the KV-cache with out re-encoding. Experiments reveal sturdy efficiency throughout picture era and multimodal understanding benchmarks, validating autoregressive flows as a viable basis for unified multimodal modeling.
- †UIUC
- ** Work completed whereas at Apple






