activity
20242026
collaborators

6 papers

cs.CV2026

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

Weixing Wang, Liudvikas Zekas, Anton Hackl +5

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these…

cs.CV2026

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

Liu Liu, Alexandra Kudaeva, Marco Cipriano +4

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such in…

cs.CL2025

Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches

Margarita Bugueño, Gerard de Melo

In document classification, graph-based models effectively capture document structure, overcoming sequence length limitations and enhancing contextual understanding. However, most…

cs.CV2024

ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection

Maryam Hosseini, Marco Cipriano, Sedigheh Eslami +4

Existing Open Vocabulary Detection (OVD) models exhibit a number of challenges. They often struggle with semantic consistency across diverse inputs, and are often sensitive to slig…

cs.CL2024

GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization

Margarita Bugueño, Hazem Abou Hamdan, Gerard de Melo

Heterogeneous graph neural networks have recently gained attention for long document summarization, modeling the extraction as a node classification task. Although effective, these…

cs.CV2024

Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision

Moritz Feuerpfeil, Marco Cipriano, Gerard de Melo

Scalable Vector Graphics (SVG) is a popular format on the web and in the design industry. However, despite the great strides made in generative modeling, SVG has remained underexpl…