When evaluating LLMs that generate hardware assertions, using equivalent implementations reveals that the same behavior can produce different numbers of correct assertions—showing that assertion quality depends on implementation details, not just the intended behavior.
EquivSVA is a dataset of 120 behavior families with 480 RTL implementations and 914 formally verified assertions, designed to test whether AI-generated hardware assertions capture true behavior or just implementation details. Each behavior family has four structurally different implementations of the same functionality, enabling controlled studies of assertion robustness.