LLM agents need specialized training on verified statistical tasks to reliably conduct hypothesis testing—standard benchmarks miss inferential errors, and reinforcement learning with statistical rewards substantially improves correctness.
LLM agents often make subtle statistical errors when conducting hypothesis testing, even when code runs correctly. This paper introduces P-Bench, a benchmark of 425 realistic hypothesis-testing tasks, and Fisher-R1, an open-weight agent trained with reinforcement learning to perform rigorous statistical analysis.