For edge AI deployment, combining hierarchical operator parallelism with pipelining can achieve better energy efficiency and latency than using either strategy alone—Para-Pipe shows 11-23% energy improvements on real SoCs.
Para-Pipe is a framework that optimizes deep learning inference on edge devices (SoCs) by combining pipelining and parallel execution of neural network operations.