Steady diffusion and circulate matching fashions might signify a strong different to autoregressive approaches for language modelling (LM), as they unlock a bunch of benefits at the moment reserved for steady modalities, together with accelerated sampling and tilting. Just lately, a number of works have demonstrated the potential of producing discrete knowledge constantly by a easy circulate matching course of between a Gaussian and the one-hot encoded knowledge distribution. They’ve additional proven the feasibility of accelerated sampling through Categorical Stream Maps (CFMs), leading to aggressive pattern high quality within the few-step regime. Nonetheless, this technique had solely been evaluated at comparatively modest scales (< 1B), leaving the query of its scalability utterly open. On this article, we practice a 1.7B-parameter base circulate mannequin on 2.1T tokens and self-distill it right into a CFM that generates various, high-quality textual content in as few as 4 inference steps whereas sustaining near-data-level token entropy. Moreover, we introduce a chance certain for CFMs within the semi-discrete setting, and present that they can be utilized to attain the mannequin on commonplace LM benchmarks, reaching leads to the identical vary as discrete diffusion strategies. Lastly, we uncover a number of the challenges that come up from coaching these fashions at scale, and we offer prescriptive insights on loss weighting and time scheduling.
- †College of Oxford
- ** Work accomplished whereas at Apple







