3 papers
cs.CV2026
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
Manuel Traub, Martin V. Butz
Current state-of-the-art segmentation models encode entire images before focusing on specific objects. This wastes computational resources. We introduce FLIP (Fovea-Like Input Patc…
cs.LG2026
Semantic Allocation in Ordered Bottlenecks: Predictive Residual Inference for Visual Representation Learning
Erik Ayari, Manuel Traub, Martin V. Butz
Ordered bottlenecks aim to provide utility at flexible budgets by assigning coarse information to early tokens and task-relevant detail to later ones. Prior work, including tail dr…
cs.CV2024
Learning Object Permanence from Videos via Latent Imaginations
Manuel Traub, Frederic Becker, Sebastian Otte +1
While human infants exhibit knowledge about object permanence from two months of age onwards, deep-learning approaches still largely fail to recognize objects' continued existence.…