7 papers · 1 filter
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
Diogo Glória-Silva, João Cardeira, Manuel Letras da Luz +8
Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which…
PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction
João Cardeira, Diogo Glória-Silva, Manuel Letras da Luz +4
European Portuguese (pt-PT) is largely absent from Optical Character Recognition (OCR) benchmarks, which skew toward high-resource languages. The few benchmarks that cover pt-PT fo…
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
Diogo Glória-Silva, David Semedo, João Maglhães
We introduce VIGiA, a novel multimodal dialogue model designed to understand and reason over complex, multi-step instructional video action plans. Unlike prior work which focuses m…
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
João Pereira, Vasco Lopes, João Neves +1
Video Anomaly Understanding (VAU) is a novel task focused on describing unusual occurrences in videos. Despite growing interest, the evaluation of VAU remains an open challenge. Ex…
Chain-of-Anomaly Thoughts with Large Vision-Language Models
Pedro Domingos, João Pereira, Vasco Lopes +2
Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning…
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
Joao Pereira, Vasco Lopes, David Semedo +1
Large Vision-Language Models (LVLMs) demonstrate remarkable performance in short-video tasks such as video question answering, but struggle in long-video understanding. The linear…