Splitting model weights across multiple GPUs so each GPU computes part of a matrix operation in parallel.