8 papers
Multimedia and Visual Analytics in the Agentic Era
Marcel Worring, Jan Zahálka, Stef van den Elzen +2
Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have rapidly changed the playing field, and i…
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
In many real-world applications, modeling both the internal structure of sets and their temporal relationships is essential for capturing complex underlying patterns. Sequential mu…
Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
Ujjwal Sharma, Stevan Rudinac, Ana MiÄkoviÄ +2
In this work, we introduce a multimodal analysis pipeline that leverages large foundation models in vision and language to analyze corporate social media content, with a focus on s…
Modeling Edge-Specific Node Features through Co-Representation Neural Hypergraph Diffusion
Yijia Zheng, Marcel Worring
Hypergraphs are widely being employed to represent complex higher-order relations in real-world applications. Most existing research on hypergraph learning focuses on node-level or…
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
Shuai Wang, Ivona Najdenkoska, Hongyi Zhu +4
Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large lan…