3 papers
cs.LG2026
Disentangling 3D Modeling from Spatial Reasoning
Haoze Sun, Jiequan Cui, Qingshan Xu +1
In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perceptio…
cs.CV2026
Visual Token Compression Enhances Robustness of MLLMs
Shishen Gu, Jiequan Cui, Wenbo Hu +3
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…
cs.LG2026
Class-frequency Guided Noise Schedule for Diffusion Models
Jiequan Cui, Beier Zhu, Qingshan Xu +3
In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, l…