Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
By Zhi Chen · Paper · cs.SE
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasing