4 papers
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
Weijie Zhao, Mingquan Liu, Bolun Wang +4
Scaling Transformers typically necessitates training larger models from scratch, as standard architectures struggle to expand without discarding learned representations. We identif…
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
Rui Zhu, Xin Shen, Shuchen Wu +6
Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmark…
FingerCap: Fine-grained Finger-level Hand Motion Captioning
Xin Shen, Rui Zhu, Lei Shen +10
Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-…
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
Yuxuan Liang, Xu Li, Xiaolei Chen +4
Large Vision-Language Models (LVLMs) have recently demonstrated strong multimodal understanding, yet their fine-grained visual perception is often constrained by low input resoluti…