WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

By Zhenran Wang · Paper · cs.CL

Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the

Model Launch · Cs.cl

View original

HomeResourceLoading…