11 papers
Locality Matters for Training-Free Audio Token Compression in Audio-Language Models
Jiale Luo, Xiaoyu Liang, Haoji Hu
Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio in…
D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation
Wenjie Zheng, Haoji Hu, Jiali Lu +2
Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image clas…
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
Jiedong Zhuang, Lu Lu, Ming Dai +4
Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual to…
Learn Before Represent: Bridging Generative and Contrastive Learning for Domain-Specific LLM Embeddings
Xiaoyu Liang, Yuchen Peng, Jiale Luo +3
Large Language Models (LLMs) adapted via contrastive learning excel in general representation learning but struggle in vertical domains like chemistry and law, primarily due to a l…
SemanticGen: Video Generation in Semantic Space
Jianhong Bai, Xiaoshi Wu, Xintao Wang +9
State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can gene…
Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting
Jiangnan Ye, Jiedong Zhuang, Lianrui Mu +5
We introduce GS-Light, an efficient, textual position-aware pipeline for text-guided relighting of 3D scenes represented via Gaussian Splatting (3DGS). GS-Light implements a traini…