QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
By Sergio Hernández-Gutiérrez · Paper · cs.LG
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision m