arXiv cs.CL
· Papers
ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation
arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML pipeline generation beyond English task des