6 papers
Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models
Weixing Wang, Liudvikas Zekas, Anton Hackl +5
Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these…
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
Liu Liu, Alexandra Kudaeva, Marco Cipriano +4
Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such in…
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
Margarita Bugueño, Gerard de Melo
In document classification, graph-based models effectively capture document structure, overcoming sequence length limitations and enhancing contextual understanding. However, most…
ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection
Maryam Hosseini, Marco Cipriano, Sedigheh Eslami +4
Existing Open Vocabulary Detection (OVD) models exhibit a number of challenges. They often struggle with semantic consistency across diverse inputs, and are often sensitive to slig…
GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization
Margarita Bugueño, Hazem Abou Hamdan, Gerard de Melo
Heterogeneous graph neural networks have recently gained attention for long document summarization, modeling the extraction as a node classification task. Although effective, these…
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision
Moritz Feuerpfeil, Marco Cipriano, Gerard de Melo
Scalable Vector Graphics (SVG) is a popular format on the web and in the design industry. However, despite the great strides made in generative modeling, SVG has remained underexpl…