Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data…