2 papers
cs.CV2026
Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models
Ruchen Liu, Yi Yang, Yiming Xu +3
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show tha…
cs.CV2026
OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation
John Helsby, Yi Yang, Bodo Rosenhahn +1
Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as subject, predicate, object triplets, and underpin downstream tasks such as video…