4 papers
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
Ioanna Ntinou, Alexandros Xenos, Yassine Ouali +2
Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these mod…
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
Alexandros Xenos, Niki Maria Foteinopoulou, Ioanna Ntinou +2
Recognising emotions in context involves identifying an individual's apparent emotions while considering contextual cues from the surrounding scene. Previous approaches to this tas…
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
Ioanna Ntinou, Enrique Sanchez, Georgios Tzimiropoulos
This paper is on long-term video understanding where the goal is to recognise human actions over long temporal windows (up to minutes long). In prior work, long temporal context is…
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
Ioanna Ntinou, Enrique Sanchez, Georgios Tzimiropoulos
Action Localization is a challenging problem that combines detection and recognition tasks, which are often addressed separately. State-of-the-art methods rely on off-the-shelf bou…