9 papers
Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference
Hui Zang, Pengfei Xia, Hong Liu +5
Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data mo…
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
Xuzhao Geng, Haozhao Wang, Xuelian Li +4
Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed…
Unbiased Rectification for Sequential Recommender Systems Under Fake Orders
Qiyu Qin, Yichen Li, Haozhao Wang +3
Fake orders pose increasing threats to sequential recommender systems by misleading recommendation results through artificially manipulated interactions, including click farming, c…
UNGER: Generative Recommendation with A Unified Code via Semantic and Collaborative Integration
Longtao Xiao, Haozhao Wang, Cheng Wang +6
With the rise of generative paradigms, generative recommendation has garnered increasing attention. The core component is the item code, generally derived by quantizing collaborati…
Resource-Constrained Federated Continual Learning: What Does Matter?
Yichen Li, Yuying Wang, Jiahua Dong +4
Federated Continual Learning (FCL) aims to enable sequentially privacy-preserving model training on streams of incoming data that vary in edge devices by preserving previous knowle…
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
Senyao Li, Haozhao Wang, Wenchao Xu +6
As large language models (LLMs) evolve, deploying them solely in the cloud or compressing them for edge devices has become inadequate due to concerns about latency, privacy, cost,…