SocietyBench: Forecasting Counterfactual Social-World Evolution
By Zhenran Wang · Paper · cs.CL
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social eve