arXiv cs.CL
· Papers
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines