LLM-based fault diagnosis for autonomous systems requires rigorous ensemble testing rather than single trials, and frontier models significantly outperform local alternatives, but success depends on models following complete diagnostic procedures rather than jumping to conclusions.
This paper presents SPAR, a simulation platform for testing how large language models can help autonomous underwater vehicles diagnose and recover from faults without human intervention.