MOSAIC: Efficient Collaborative Inference for Multi-Modal Large Language Models with Distributed Edge Computing

PACE Overview

Abstract

Multi-modal large language models (MLLMs) enable natural language reasoning over heterogeneous sensory data, appealing to many edge AI scenarios, yet their deployment faces competing demands on latency, energy, and quality. We present MOSAIC, which adopts a component-level decomposition and distributed deployment paradigm for efficient MLLM inference at the edge, realizing three core ideas parallel encoding, token pruning, and device-server collaboration. Specifically, for the offline stage, modality-specific encoders distilled from a unified foundation model are co-deployed with Q-Former projection on edge devices for parallel encoding, while LoRA instruction fine-tuning adapts the frozen LLM backbone to downstream tasks. For the online stage, we formulate a multiobjective optimization over latency and energy that jointly tunes token budgets, device frequencies, and quantization precision under long-term semantic accuracy constraints, achieving dynamic token pruning and device-server collaboration. We leverage Lagrangian relaxation to decompose the original problem into tractable per-slot subproblems and design an online algorithm with provable nearoptimal performance and guaranteed constraint satisfaction. We implement a realistic testbed for extensive performance evaluation, which demonstrates that MOSAIC reduces latency by up to 49.8% and energy by up to 65.1% over advanced baselines while maintaining high semantic accuracy.

Publication
In International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing (MOBIHOC), Tokyo, Japan, 23-26 November 2026, CCF-B, Acceptance rate =23%.
Shengyuan Ye
Shengyuan Ye
Ph.D. in Computer Science

Shengyuan Ye obtained his Ph.D. degree from the School of Computer Science and Engineering, Sun Yat‑sen University. His research interests include AI for Power System and Power Dispatching Automation.

Liekang Zeng
Liekang Zeng
Ph.D., Sun Yat-sen University

He obtained Ph.D. degree at School of Computer Science and Engineering, Sun Yat-sen University. His research interest lies in building edge intelligence systems with real-time responsiveness, systematic resource efficiency, and theoretical performance guarantee.

Xu Chen
Xu Chen
Professor and Assistant Dean, Sun Yat-sen University
Director, Institute of Advanced Networking & Computing Systems

Xu Chen is a Full Professor with Sun Yat-sen University, Director of Institute of Advanced Networking and Computing Systems (IANCS), and the Vice Director of National Engineering Research Laboratory of Digital Homes. His research interest includes edge computing and cloud computing, federated learning, cloud-native intelligent robots, distributed artificial intelligence, intelligent big data analysis, and computing power network.