Distributed RL agents that share only gradients—not raw data—still leak sensitive trajectory information through temporal correlations; defending against this requires sequence-aware privacy mechanisms, not just per-step protections.
This paper reveals a critical privacy vulnerability in distributed embodied AI systems. When agents send policy gradients to a server instead of raw sensor data, attackers can reconstruct the agent's complete trajectory of observations and actions by analyzing the temporal patterns in these gradients.