4 papers
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Jincheng Xie, Runheng Liu, Heyan Huang +4
Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activat…
AdaPLD: Adaptive Retrieval and Reuse for Efficient Model-Free Speculative Decoding
Runheng Liu, Jincheng Xie, Wen Hu +2
Speculative decoding accelerates generation by verifying multiple drafted tokens in a single target-model forward pass, reducing sequential decoding iterations. Model-free variants…
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
Xingchen Xiao, Heyan Huang, Runheng Liu +1
Large language models (LLMs) are widely used in retrieval-augmented generation (RAG) to incorporate external knowledge at inference time. However, when retrieved contexts are noisy…
Joint single-shot ToA and DoA estimation for VAA-based BLE ranging with phase ambiguity: A deep learning-based approach
Jincheng Xie, Yili Deng, Jiguang He +4
Conventional direction-of-arrival (DoA) estimation methods rely on multi-antenna arrays, which are costly to implement on size-constrained Bluetooth Low Energy (BLE) devices. Virtu…