Skip to content
HF Daily Papers · Papers

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two comple