QuoteBench: How Matched Scores Can Hide Command-Path Failures
By Shangao Li · Paper · cs.AI
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-s