118 citations · 423 across the 32 of their papers we have counts for
7 papers · 1 filter
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
Zongzhao Li, Xiangzhe Kong, Jiahui Su +8
This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entitie…
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Ruifeng Yuan, Chenghao Xiao, Sicong Leng +9
Reinforcement learning has proven its effectiveness in enhancing the reasoning capabilities of large language models. Recent research efforts have progressively extended this parad…
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
Zongzhao Li, Zongyang Ma, Mingze Li +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, yet they lag significantly behind humans in spatial reasoning. We investiga…
Towards Complete-View and High-Level Pose-based Gait Recognition
Honghu Pan, Yongyong Chen, Tingyang Xu +2
The model-based gait recognition methods usually adopt the pedestrian walking postures to identify human beings. However, existing methods did not explicitly resolve the large intr…
Smoothing Matters: Momentum Transformer for Domain Adaptive Semantic Segmentation
Runfa Chen, Yu Rong, Shangmin Guo +4
After the great success of Vision Transformer variants (ViTs) in computer vision, it has also demonstrated great potential in domain adaptive semantic segmentation. Unfortunately,…
Deep Multimodal Fusion by Channel Exchanging
Yikai Wang, Wenbing Huang, Fuchun Sun +3
Deep multimodal fusion by using multiple sources of data for classification or regression has exhibited a clear advantage over the unimodal counterpart on various applications. Yet…