Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design — ThinkLLM