Skip to content
arXiv cs.CL · Papers

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML pipeline generation beyond English task des