Score: 2

OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference

Published: December 3, 2025 | arXiv ID: 2512.03927v1

By: Liujianfu Wang , Yuyang Du , Yuchen Pan and more

Potential Business Impact:

Lets small computers run big AI models.

Business Areas:

Cloud Computing Internet Services, Software

Mixture-of-Experts (MoE), while offering significant advantages as a Large Language Model (LLM) architecture, faces substantial challenges when deployed on low-cost edge devices with tight memory constraints. Expert offloading mitigates this issue by storing expert parameters in CPU memory and caching a subset of popular experts in GPU memory. Although this approach improves GPU memory utilization by caching only the likely-used experts, the GPU memory reserved for expert caching is underutilized compared with dense LLMs. This paper presents OD-MoE, a distributed MoE inference framework that obviates the need for expert caches via fully on-demand expert loading. OD-MoE is built upon two key mechanisms: 1) parallelizing expert loading and expert computation across distributed edge nodes, and 2) an ultra-accurate emulative predictor that forecasts expert activations multiple layers ahead while expert computation is ongoing. With these innovations, OD-MoE dynamically loads each target expert to one of the distributed nodes just-in-time before its activation and promptly evicts it afterward, freeing GPU memory for subsequent experts. We comprehensively benchmark OD-MoE against state-of-the-art MoE offloading systems on a ten-node testbed. Experimental results show that: 1) OD-MoE achieves 99.94% expert activation prediction accuracy, substantially surpassing all existing methods; and 2) OD-MoE delivers approximately 75% of the decoding speed of a fully GPU-cached MoE deployment while using only 1/3 of the GPU memory. More importantly, by eliminating the need for expert caches, OD-MoE enables MoE inference on edge nodes with less-than-1GB GPU memory, paving the way for practical MoE deployment of low-cost IoT devices at the edge in the LLM era.

CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge

Networking and Internet Architecture

Makes big AI models fit on phones.

10 Aug 2025 1

93%

BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference

Machine Learning (CS)

Lets AI learn more without needing more computer memory.

13 Nov 2025 1

93%

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Distributed, Parallel, and Cluster Computing

Lets smart computers share tasks on many devices.

18 Aug 2025 2

View PDF Login to Bookmark

Repos / Data Links

github.com

Page Count

14 pages

OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference

Lets small computers run big AI models.

Technical Abstract

CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge

BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement