LLM-generated test harnesses for safety-critical autonomous vehicle code fail primarily due to build system complexity, not reasoning ability—suggesting that tooling and integration matter more than model capability for dynamic vulnerability analysis.
This paper investigates whether LLMs can automatically generate executable test artifacts to confirm exploitable weaknesses in Autoware, an autonomous vehicle software stack.