Skip to content
arXiv cs.CL · Papers

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures

arXiv:2607.02266v1 Announce Type: cross Abstract: Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer can express. Existing labels, including provenance, topic or format taxonomies, and flat embedding clusters, commit to one semanti