collaborators

6 papers

cs.CV2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

Qihui Zhu, Tao Zhang, Yuchen Wang +9

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time…

cs.AI2026

Improving Safety Alignment via Balanced Direct Preference Optimization

Shiji Zhao, Mengyang Wang, Shukun Xiong +7

With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from…

cs.AI2026

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

Qihui Zhu, Shouwei Ruan, Xiao Yang +6

Despite the widespread adoption of MLLMs in embodied agents, their capabilities remain largely confined to reactive planning from immediate observations, consistently failing in sp…

cs.AI2026

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

Shouwei Ruan, Bin Wang, Zhenyu Wu +4

Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D gro…

cs.AI2025

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

Shouwei Ruan, Liyuan Wang, Caixin Kang +4

Spatial cognition enables adaptive goal-directed behavior by constructing internal models of space. Robust biological systems consolidate spatial knowledge into three interconnecte…

cs.CV2025

MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models

Shiji Zhao, Qihui Zhu, Shukun Xiong +7

Large pre-trained Vision Language Models (VLMs) demonstrate excellent generalization capabilities but remain highly susceptible to adversarial examples, posing potential security r…