13 papers · 1 filter
OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation
John Helsby, Yi Yang, Bodo Rosenhahn +1
Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as subject, predicate, object triplets, and underpin downstream tasks such as video…
Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models
Ruchen Liu, Yi Yang, Yiming Xu +3
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show tha…
PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation
Yi Yang, Myrna Castillo, Bodo Rosenhahn +1
Online 3D scene graph generation builds a persistent, structured representation of a scene by incrementally fusing 2D observations into a global 3D graph. Existing online methods t…
MATCH: Flow Matching for Multi-View Anomaly Detection
Mathis Kruse, Melissa Schween, Bodo Rosenhahn
Detecting anomalies in industrial objects is an important topic for increasing production efficiency. More complex objects often require the analysis of several view points, which…
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
Robert Zimmermann, Thomas Norrenbrock, Bodo Rosenhahn
Although visual foundation models like DINOv2 provide state-of-the-art performance as feature extractors, their complex, high-dimensional representations create substantial hurdles…
Video Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction
Patrick Glandorf, Thomas Norrenbrock, Bodo Rosenhahn
Vision Transformers (ViTs) have demonstrated state-ofthe-art performance in several benchmarks, yet their high computational costs hinders their practical deployment. Patch Pruning…