Distributed inference pipelines that merge fast and slow predictions are vulnerable to network-based denial-of-deadline attacks that don't require model access—attackers can degrade accuracy by simply congesting shared resources to miss latency deadlines.
This paper reveals a new vulnerability in distributed inference systems that combine fast local predictions with slower, more accurate remote predictions. Attackers can craft burst-shaped workloads that congest the network, delaying remote predictions past their deadlines so they get discarded.