3 papers
cs.CV2026
MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
Lulu Hu, Wenhu Xiao, Xin Chen +4
Post-training quantization (PTQ) with computational invariance for Large Language Models~(LLMs) have demonstrated remarkable advances, however, their application to Multimodal Larg…
cs.CL2026
D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use
Bowen Xu, Shaoyu Wu, Hao Jiang +4
Effective tool use and reasoning are essential capabilities for large reasoning models~(LRMs) to address complex real-world problems. Through empirical analysis, we identify that c…
cs.CL2025
La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
Kai Liu, Bowen Xu, Shaoyu Wu +4
Activation sparsity can reduce the computational overhead and memory transfers during the forward pass of Large Language Model (LLM) inference. Existing methods face limitations, e…