activity
20242026
collaborators

10 papers

cs.CV2026

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Yuzhang Shang, Mu Cai, Bingxin Xu +2

Large Multimodal Models (LMMs) have shown significant visual reasoning capabilities by connecting a visual encoder and a large language model. LMMs typically take in a fixed and la…

cs.RO2026

Real-Time Robot Execution with Masked Action Chunking

Haoxuan Wang, Gengyu Zhang, Yan Yan +3

Real-time execution is essential for cyber-physical systems such as robots. These systems operate in dynamic real-world environments where even small delays can undermine responsiv…

cs.CV2025

Distill Video Datasets into Images

Zhenghao Zhao, Haoxuan Wang, Kai Wang +3

Dataset distillation aims to synthesize compact yet informative datasets that allow models trained on them to achieve performance comparable to training on the full dataset. While…

cs.CV2025

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu +3

Dataset distillation aims to synthesize a small dataset from a large dataset, enabling the model trained on it to perform well on the original dataset. With the blooming of large l…

cs.CL2025

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Zhihang Yuan, Yuzhang Shang, Yue Song +4

In this paper, we introduce a new post-training compression paradigm for Large Language Models (LLMs) to facilitate their wider adoption. We delve into LLM weight low-rank decompos…

cs.CV2025

QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning

Haoxuan Wang, Yuzhang Shang, Zhihang Yuan +3

The practical deployment of diffusion models is still hindered by the high memory and computational overhead. Although quantization paves a way for model compression and accelerati…