Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

By Jiacheng Miao · Paper · cs.AI

Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently

Cs.ai

View original

HomeResourceLoading…