SABRE: Scalable and Automated Benchmarking of VLMs under Stress
By Zixuan Lan · Paper · cs.CV
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalab