6 papers
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki +2
One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectiv…
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2
Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large language models (MLLMs). Ho…
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2
Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they still struggle with object hall…
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
Zilong Deng, Federico Tombari, Marc Pollefeys +2
Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the n…
Towards Real-Time Open-Vocabulary Video Instance Segmentation
Bin Yan, Martin Sundermeyer, David Joseph Tan +2
In this paper, we address the challenge of performing open-vocabulary video instance segmentation (OV-VIS) in real-time. We analyze the computational bottlenecks of state-of-the-ar…
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
Fan-Yun Sun, Weiyu Liu, Siyi Gu +6
Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demon…