9 papers
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
Liulu He, XuanAng Liu, Juntao Liu +8
Existing quantization methods are fundamentally limited by rigid, integer-based bit-widths (e.g., 2, 3-bit), resulting in a ``deployment gap" where Large Language Models cannot be…
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
Chong Cheng, Peilin Tao, Nanjie Yao +9
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or…
MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation
Ronyu Zhang, Aosong Cheng, Gaole Dai +8
Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture-biased backbones risk error…
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
Aojie Jiang, Kang Zhu, Zhiheng Zhang +4
Tensor parallelism (TP) has become a key technique for latency-sensitive LLM inference, but it introduces frequent, tightly synchronized All-Reduce operations that lie directly on…
Key-Embedded Privacy for Decentralized AI in Biomedical Omics
Rongyu Zhang, Hongyu Dong, Gaole Dai +13
The rapid adoption of data-driven methods in biomedicine has intensified concerns over privacy, governance, and regulation, limiting raw data sharing and hindering the assembly of…
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
Hengyu Fang, Yijiang Liu, Yuan Du +2
Vision-Language-Action (VLA) models exhibit unprecedented capabilities for embodied intelligence. However, their extensive computational and memory costs hinder their practical dep…