Encrypted reasoning traces from LLM APIs are architecturally vulnerable to cross-model decryption attacks—adversaries can extract proprietary reasoning by exploiting compatibility between encrypted blocks across different models in the same provider's ecosystem.
Researchers discovered that major LLM providers (OpenAI, Anthropic, Google) encrypt reasoning traces sent to clients in a way that makes them reusable across different sessions and models. By injecting encrypted reasoning from one model into a weaker model, attackers can force it to decrypt and reveal the reasoning in plaintext.