Designing algorithms to minimize data movement between GPU memory and compute units, addressing the bottleneck of memory bandwidth rather than just computation.