4 papers
MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
Lulu Hu, Wenhu Xiao, Xin Chen +4
Post-training quantization (PTQ) with computational invariance for Large Language Models~(LLMs) have demonstrated remarkable advances, however, their application to Multimodal Larg…
D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use
Bowen Xu, Shaoyu Wu, Hao Jiang +4
Effective tool use and reasoning are essential capabilities for large reasoning models~(LRMs) to address complex real-world problems. Through empirical analysis, we identify that c…
La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
Kai Liu, Bowen Xu, Shaoyu Wu +4
Activation sparsity can reduce the computational overhead and memory transfers during the forward pass of Large Language Model (LLM) inference. Existing methods face limitations, e…
Mixture-of-Instructions: Aligning Large Language Models via Mixture Prompting
Bowen Xu, Shaoyu Wu, Kai Liu +1
With the proliferation of large language models (LLMs), the comprehensive alignment of such models across multiple tasks has emerged as a critical area of research. Existing alignm…