1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
In many real-world applications, modeling both the internal structure of sets and their temporal relationships is essential for capturing complex underlying patterns. Sequential mu…
Graph Neural Networks for Knowledge Enhanced Visual Representation of Paintings
Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic +2
We propose ArtSAGENet, a novel multimodal architecture that integrates Graph Neural Networks (GNNs) and Convolutional Neural Networks (CNNs), to jointly learn visual and semantic-b…
TULIP: Token-length Upgraded CLIP
Ivona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki M. Asano +3
We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restrict…