Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.
EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.