8 papers
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
Aojie Jiang, Kang Zhu, Zhiheng Zhang +4
Tensor parallelism (TP) has become a key technique for latency-sensitive LLM inference, but it introduces frequent, tightly synchronized All-Reduce operations that lie directly on…
Key-Embedded Privacy for Decentralized AI in Biomedical Omics
Rongyu Zhang, Hongyu Dong, Gaole Dai +13
The rapid adoption of data-driven methods in biomedicine has intensified concerns over privacy, governance, and regulation, limiting raw data sharing and hindering the assembly of…
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
Qi Wu, Chao Fang, Jiayuan Chen +5
Mixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems wi…
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
Rongyu Zhang, Jiaming Liu, Xiaoqi Li +5
Vision-centric Bird's Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issu…
RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
Rongyu Zhang, Xize Duan, Jiaming Liu +5
Recently, content-aware methods have been employed to reduce bandwidth and enhance the quality of Internet video delivery. These methods involve training distinct content-aware sup…
FBQuant: FeedBack Quantization for Large Language Models
Yijiang Liu, Hengyu Fang, Liulu He +4
Deploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user p…