From the 1 of 11 linked papers with an AI index.
4 papers · 1 filter
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, Xiaotian Zhang +4
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduct…
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding
Yifan Wang, Peiming Li, Shiyu Li +5
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Ziyi Wang, Peiming Li, Xinshun Wang +3
Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skeletons. Existing methods either c…
SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
Yifan Wang, Yian Zhao, Fanqi Pu +4
Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to esti…