Every large language model begins as a question of proportion. How much web text? How much code? How much scientific literature? The ratio of training data across these domains — what researchers call ...