Appearance pointers provide a modality-agnostic way to add regional control to existing Diffusion Transformers, letting you specify exactly where text descriptions or image references should apply in the generated output without expensive retraining.
This paper introduces appearance pointers, a new technique for controlling where and how text and image inputs influence image generation in Diffusion Transformers.