TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
By Jiarui Yan · Paper · cs.LG
Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human co