2.7k citations · 3k across the 7 of their papers we have counts for
5 papers · 1 filter
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…
Pay attention to the activations: a modular attention mechanism for fine-grained image recognition
Pau Rodríguez López, Diego Velazquez Dorta, Guillem Cucurull Preixens +3
Fine-grained image recognition is central to many multimedia tasks such as search, retrieval and captioning. Unfortunately, these tasks are still challenging since the appearance o…
Context-Aware Visual Compatibility Prediction
Guillem Cucurull, Perouz Taslakian, David Vazquez
How do we determine whether two or more clothing items are compatible or visually appealing? Part of the answer lies in understanding of visual aesthetics, and is biased by persona…
Attend and Rectify: a Gated Attention Mechanism for Fine-Grained Recovery
Pau Rodríguez, Josep M. Gonfaus, Guillem Cucurull +2
We propose a novel attention mechanism to enhance Convolutional Neural Networks for fine-grained recognition. It learns to attend to lower-level feature activations without requiri…
On the iterative refinement of densely connected representation levels for semantic segmentation
Arantxa Casanova, Guillem Cucurull, Michal Drozdzal +2
State-of-the-art semantic segmentation approaches increase the receptive field of their models by using either a downsampling path composed of poolings/strided convolutions or succ…