You can improve a frozen language model's reasoning and confidence detection in one forward pass by reconstructing clean internal states after steering, rather than running separate passes or accepting interference between techniques.
This paper solves a key problem in frozen language models: they both misuse their internal knowledge and fail to recognize when they lack sufficient information to answer.