4 papers
DeepLatent: Think with Images via Parallel Latent Visual Reasoning
Dongchen Lu, Zhimo Li, Mao Shu +1
The emerging paradigm of "thinking with images" embeds visual states into intermediate reasoning steps, defining a new frontier for Vision-Language Models. Existing approaches dive…
RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
Xinhua Wang, Kun Wu, Zhen Zhao +10
Enhancing the generalization capability of robotic learning to enable robots to operate effectively in diverse, unseen scenes is a fundamental and challenging problem. Existing app…
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Dongchen Lu, Yuyao Sun, Zilu Zhang +4
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great qua…
UniLoc: Towards Universal Place Recognition Using Any Single Modality
Yan Xia, Zhendong Li, Yun-Jin Li +4
To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allow…