Generated text can carry cryptographically verifiable evidence about which internal computations a model actually performed, opening possibilities for auditing model reasoning without changing the visible output.
This paper demonstrates that language models can embed hidden evidence of their internal computational states into generated text. Researchers trained neural networks on arithmetic tasks with mandatory intermediate decision points, then verified which internal state was used and encoded that information as subtle statistical patterns in the output.