9 papers
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
Guanqiao Qu, Shuo Chen, Qian Chen +2
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: A…
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Oscar Chew, Serhii Honcharenko, Qian-Hui Chen +4
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains uncle…
TubiFM: Unified Item, Carousel, and Search Ranking for Streaming Discovery
Alexandre Salle, Chenglei Niu, Suchismit Mahapatra +7
Personalized discovery systems often train separate models for item ranking, carousel ranking, and search, even though these tasks expose complementary signals from the same viewer…
SpaceMoE: Towards Orbital General Intelligence with Distributed Mixture-of-Experts Inference
Qian Chen, Xianhao Chen, Min Sheng +1
As satellite networks evolve to support increasingly diverse services and artificial general intelligence (AGI), large language models (LLMs) are emerging as a critical foundation…
SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
Guanqiao Qu, Tao Li, Qian Chen +2
To support on-device inference, the next-generation mobile networks are expected to support real-time model downloading services to mobile users. However, powerful AI models typica…
SiftMoE: Similarity-Aware Energy-Efficient Expert Selection for Wireless Distributed MoE Inference
Qian Chen, Xianhao Chen, Kaibin Huang
Mixture-of-Experts (MoE) architectures leverage sparse activation to enhance the scalability of large language models (LLMs), making them suitable for deployment in resource-constr…