HF Daily Papers
· Papers
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data Mixing (ODM), the technique of adaptively adjusting data mixtures during training, has emerged as a promising direction to improv