papers

Publications (23)

cs.CV2024

Towards Real-Time Open-Vocabulary Video Instance Segmentation

Bin Yan, Martin Sundermeyer, David Joseph Tan +2

In this paper, we address the challenge of performing open-vocabulary video instance segmentation (OV-VIS) in real-time. We analyze the computational bottlenecks of state-of-the-ar…

cs.CV2020

Self-Supervised Object-in-Gripper Segmentation from Robotic Motions

Wout Boerdijk, Martin Sundermeyer, Maximilian Durner +1

Accurate object segmentation is a crucial task in the context of robotic manipulation. However, creating sufficient annotated training data for neural networks is particularly time…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2022

Iterative Corresponding Geometry: Fusing Region and Depth for Highly Efficient 3D Tracking of Textureless Objects

Manuel Stoiber, Martin Sundermeyer, Rudolph Triebel

Tracking objects in 3D space and predicting their 6DoF pose is an essential task in computer vision. State-of-the-art approaches often rely on object texture to tackle this problem…

cs.CV2025

BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation

Van Nguyen Nguyen, Stephen Tyree, Andrew Guo +16

We present the evaluation methodology, datasets and results of the BOP Challenge 2024, the 6th in a series of public competitions organized to capture the state of the art in 6D ob…

cs.CV2023

BOP Challenge 2022 on Detection, Segmentation and Pose Estimation of Specific Rigid Objects

Martin Sundermeyer, Tomas Hodan, Yann Labbe +5

We present the evaluation methodology, datasets and results of the BOP Challenge 2022, the fourth in a series of public competitions organized with the goal to capture the status q…

cs.RO2021

Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes

Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel +1

Grasping unseen objects in unconstrained, cluttered environments is an essential skill for autonomous robotic manipulation. Despite recent progress in full 6-DoF grasp learning, ex…

cs.CV2020

BOP Challenge 2020 on 6D Object Localization

Tomas Hodan, Martin Sundermeyer, Bertram Drost +5

This paper presents the evaluation methodology, datasets, and results of the BOP Challenge 2020, the third in a series of public competitions organized with the goal to capture the…

cs.CV2019

BlenderProc

Maximilian Denninger, Martin Sundermeyer, Dominik Winkelbauer +5

BlenderProc is a modular procedural pipeline, which helps in generating real looking images for the training of convolutional neural networks. These can be used in a variety of use…

cs.CV2026

Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

Tom Fischer, Martin Sundermeyer, Adam Kortylewski +1

Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing such frames has traditionally…

cs.CV2019

Implicit 3D Orientation Learning for 6D Object Detection from RGB Images

Martin Sundermeyer, Zoltan-Csaba Marton, Maximilian Durner +2

We propose a real-time RGB-based pipeline for object detection and 6D pose estimation. Our novel 3D orientation estimation is based on a variant of the Denoising Autoencoder that i…

cs.CV2026

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

Junwen Huang, Jiaqi Hu, Peter KT Yu +3

While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical complexities of industrial envi…

cs.CV2026

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki +2

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectiv…

cs.CV2026

TAPNext++: What's Next for Tracking Any Point (TAP)?

Sebastian Jung, Artem Zholus, Martin Sundermeyer +6

Tracking-Any-Point (TAP) models aim to track any point through a video which is a crucial task in AR/XR and robotics applications. The recently introduced TAPNext approach proposes…

cs.CV2021

Unknown Object Segmentation from Stereo Images

Maximilian Durner, Wout Boerdijk, Martin Sundermeyer +3

Although instance-aware perception is a key prerequisite for many autonomous robotic applications, most of the methods only partially solve the problem by focusing solely on known…

cs.CV2024

BOP Challenge 2023 on Detection, Segmentation and Pose Estimation of Seen and Unseen Rigid Objects

Tomas Hodan, Martin Sundermeyer, Yann Labbe +7

We present the evaluation methodology, datasets and results of the BOP Challenge 2023, the fifth in a series of public competitions organized to capture the state of the art in mod…

cs.CV2024

HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation

Yongliang Lin, Yongzhi Su, Praveen Nathan +7

In this work, we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive p…

cs.CV2023

6D Object Pose Estimation from Approximate 3D Models for Orbital Robotics

Maximilian Ulmer, Maximilian Durner, Martin Sundermeyer +2

We present a novel technique to estimate the 6D pose of objects from single images where the 3D geometry of the object is only given approximately and not as a precise 3D model. To…

cs.CV2023

A Multi-body Tracking Framework - From Rigid Objects to Kinematic Structures

Manuel Stoiber, Martin Sundermeyer, Wout Boerdijk +1

Kinematic structures are very common in the real world. They range from simple articulated objects to complex mechanical systems. However, despite their relevance, most model-based…

cs.CV2026

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

Luca Barsellotti, Martin Sundermeyer, Mattia Segu +5

Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computation…

cs.CV2020

Multi-path Learning for Object Pose Estimation Across Domains

Martin Sundermeyer, Maximilian Durner, En Yen Puang +4

We introduce a scalable approach for object pose estimation trained on simulated RGB views of multiple 3D models together. We learn an encoding of object views that does not only d…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CV2021

"What's This?" -- Learning to Segment Unknown Objects from Manipulation Sequences

Wout Boerdijk, Martin Sundermeyer, Maximilian Durner +1

We present a novel framework for self-supervised grasped object segmentation with a robotic manipulator. Our method successively learns an agnostic foreground segmentation followed…