activity
20242026
collaborators

7 papers

cs.LG2026

Training-Free Hashing-Based Attention via Binary Principal Components

Daohai Yu, Zhanpeng Zeng, Keyu Chen +6

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decodi…

cs.LG2026

Empirical Bayes Conformal Prediction for Vision and Language Models

Jiapeng Zeng, Yogesh Prabhu, Zhanpeng Zeng +2

Conformal prediction (CP) gives distribution-free coverage for modern vision and language models, but it is often forced to make a ranking decision from a single unstable nonconfor…

cs.LG2026

Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation

Luxi Lin, Zhihang Lin, Zhanpeng Zeng +5

Speculative decoding accelerates LLM inference but suffers from performance degradation when target models are fine-tuned for specific domains. A naive solution is to retrain draft…

cs.CV2025

Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers

Yunshan Zhong, Yuyao Zhou, Yuxin Zhang +5

Data-free quantization (DFQ) enables model quantization without accessing real data, addressing concerns regarding data security and privacy. With the growing adoption of Vision Tr…

cs.CV2025

Speculative Decoding Reimagined for Multimodal Large Language Models

Luxi Lin, Zhihang Lin, Zhanpeng Zeng +1

This paper introduces Multimodal Speculative Decoding (MSD) to accelerate Multimodal Large Language Models (MLLMs) inference. Speculative decoding has been shown to accelerate Larg…

cs.CV2025

LightMotion: A Light and Tuning-free Method for Simulating Camera Motion in Video Generation

Quanjian Song, Zhihang Lin, Zhanpeng Zeng +3

Existing camera motion-controlled video generation methods face computational bottlenecks in fine-tuning and inference. This paper proposes LightMotion, a light and tuning-free met…