1 citations · 2 across the 4 of their papers we have counts for
3 papers · 1 filter
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding
Shuai Wang, Wangyuan Ding, Yixian Shen +5
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
In many real-world applications, modeling both the internal structure of sets and their temporal relationships is essential for capturing complex underlying patterns. Sequential mu…
Graph Neural Networks for Knowledge Enhanced Visual Representation of Paintings
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
We propose ArtSAGENet, a novel multimodal architecture that integrates Graph Neural Networks (GNNs) and Convolutional Neural Networks (CNNs), to jointly learn visual and semantic-b…