As language fashions scale, the quantity of information they require grows – but many goal knowledge sources, akin to low-resource languages or specialised domains, are inherently restricted in dimension. A typical technique is to combine this scarce however beneficial goal knowledge with considerable generic knowledge, which presents a basic trade-off: too little goal knowledge within the combination underexposes the mannequin to the goal area, whereas an excessive amount of goal knowledge repeats the identical examples excessively, yielding diminishing returns and eventual overfitting. We research this trade-off throughout greater than 2,000 language-model coaching runs spanning a number of mannequin and goal dataset sizes, in addition to a number of knowledge sorts, together with multilingual, domain-specific, and quality-filtered mixtures. Throughout all settings, we discover that repetition is a central driver of target-domain efficiency, and that combination coaching tolerates a lot greater repetition than single-source coaching: scarce goal corpora will be reused 15–20 occasions, with the optimum variety of repetitions relying on the goal knowledge dimension, compute funds, and mannequin scale. Subsequent, we introduce a repetition-aware combination scaling legislation that accounts for the reducing worth of repeated goal tokens and the regularizing function of generic knowledge. Optimizing the scaling legislation offers a principled approach to compute efficient combination configurations, yielding sensible combination suggestions for pretraining underneath knowledge constraints.






