4 papers
EarlyTom: Early Token Compression Completes Fast Video Understanding
Hesong Wang, Xin Jin, Lu Lu +4
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficien…
ERC-SVD: Error-Controlled SVD for Large Language Model Compression
Haolei Bai, Siyong Jian, Tuo Liang +2
Large language models (LLMs) have demonstrated impressive capabilities in a wide range of downstream natural language processing tasks. Nevertheless, their considerable sizes and m…
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
Xin Jin, Siyuan Li, Siyong Jian +2
Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language mo…
SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
Siyong Jian, Huan Wang
Autoregressive image generation models like Janus-Pro produce high-quality images, but at the significant cost of high memory and ever-growing computational demands due to the larg…