WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
By Zhenran Wang · Paper · cs.CL
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the