activity
20242026
collaborators

7 papers

cs.CV2026

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning

Hengbo Xu, Shengjie Jin, Yanbiao Ma +1

With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual…

cs.AI2026

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

Yanqi Dai, Yuxiang Ji, Xiao Zhang +3

Reinforcement Learning with Verifiable Rewards (RLVR) offers a robust mechanism for enhancing mathematical reasoning in large models. However, we identify a systematic lack of emph…

cs.AI2026

Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty

Yanqi Dai, Yong Wang, Zebin You +3

Visual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal…

cs.AI2025

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Weiliang Tang, Dong Jing, Jia-Hui Pan +5

Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. O…

cs.AI2025

Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions

Dong Jing, Nanyi Fei, Zhiwu Lu

In the realm of Large Multi-modal Models (LMMs), the instruction quality during the visual instruction tuning stage significantly influences the performance of modality alignment.…

cs.AI2025

MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents

Yanqi Dai, Huanran Hu, Lei Wang +3

Recently, Role-Playing Agents (RPAs) have garnered increasing attention for their potential to deliver emotional value and facilitate sociological research. However, existing studi…