Improving LLM reasoning without external validators like compilers or trained reward models, relying only on the model itself.