3 papers
cs.LG2026
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
Xin Zhang, Qiqi Tao, Jiawei Du +2
Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without…
cs.CV2026
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
Lei Yao, Yong Chen, Yuejiao Su +3
Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. In…
cs.CV2024
T-Mamba: A unified framework with Long-Range Dependency in dual-domain for 2D & 3D Tooth Segmentation
Jing Hao, Yonghui Zhu, Lei He +3
Tooth segmentation is a pivotal step in modern digital dentistry, essential for applications across orthodontic diagnosis and treatment planning. Despite its importance, this proce…