arXiv cs.CL
· Papers
SocietyBench: Forecasting Counterfactual Social-World Evolution
arXiv:2608.04009v2 Announce Type: replace Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way rea