3 papers
cs.CV2026
TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
Alexandros Stergiou
How do video understanding models acquire their answers? Although current Vision Language Models (VLMs) reason over complex scenes with diverse objects, action performances, and sc…
cs.CV2026
MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
Animesh Jain, Alexandros Stergiou
Vision Language Models (VLMs) encode multimodal inputs over large, complex, and difficult-to-interpret architectures, which limit transparency and trust. We propose a Multimodal In…
cs.CV2024
LAVIB: A Large-scale Video Interpolation Benchmark
Alexandros Stergiou
This paper introduces a LArge-scale Video Interpolation Benchmark (LAVIB) for the low-level video task of Video Frame Interpolation (VFI). LAVIB comprises a large collection of hig…