Standard circuit evaluation metrics can systematically prefer incorrect mechanisms, making better discovery algorithms insufficient without fixing the evaluation objective itself.
This paper reveals a critical flaw in how mechanistic interpretability evaluates circuit discovery: the standard faithfulness metrics can prefer worse circuits that merely reproduce model behavior without actually capturing the underlying mechanisms.