14 papers
Towards Metric-Agnostic Trajectory Forecasting
Markus Knoche, Daan de Geus, Bastian Leibe
Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. W…
Surgical Anatomy Recognition with Context Learning using Foundation Representations
Ronald L. P. D. de Jong, Tim J. M. Jaspers, Raf A. H. Vervoort +9
Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surgical computer vision due to…
SurGe: Improved Surface Geometry in Point Maps
Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt +4
Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well. However, their predictions still exhibit inaccurate local surface g…
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
Kadir Yilmaz, Adrian Kruse, Tristan Höfer +2
Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This keeps the field…
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Tommie Kerssies, Gabriele Berton, Ju He +5
Anticipating diverse future states is a central challenge in video world modeling. Discriminative world models produce a deterministic prediction that implicitly averages over poss…
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
Niccolò Cavagnero, Narges Norouzi, Gijs Dubbelman +1
Vision Foundation Models (VFMs) pre-trained at scale enable a single frozen encoder to serve multiple downstream tasks simultaneously. Recent VFM-based encoder-only models for imag…