QuoteBench: How Matched Scores Can Hide Command-Path Failures

By Shangao Li · Paper · cs.AI

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-s

Cs.ai

View original

HomeResourceLoading…