14 papers
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding
Shuai Wang, Wangyuan Ding, Yixian Shen +5
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
Gowreesh Mago, Pascal Mettes, Stevan Rudinac
The automatic understanding of video content is advancing rapidly. Empowered by deeper neural networks and large datasets, machines are increasingly capable of understanding what i…
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu +5
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker id…
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
In many real-world applications, modeling both the internal structure of sets and their temporal relationships is essential for capturing complex underlying patterns. Sequential mu…
VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KG…