papers

Publications (53)

cs.CV2021

Subpixel Heatmap Regression for Facial Landmark Localization

Adrian Bulat, Enrique Sanchez, Georgios Tzimiropoulos

Deep Learning models based on heatmap regression have revolutionized the task of facial landmark localization with existing models working robustly under large poses, non-uniform i…

cs.CV2022

REST: REtrieve & Self-Train for generative action recognition

Adrian Bulat, Enrique Sanchez, Brais Martinez +1

This work is on training a generative action/video recognition model whose output is a free-form action-specific caption describing the video (rather than an action class label). A…

cs.CV2026

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas +2

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates…

cs.LG2020

Tensor Dropout for Robust Learning

Arinbjörn Kolbeinsson, Jean Kossaifi, Yannis Panagakis +4

CNNs achieve remarkable performance by leveraging deep, over-parametrized architectures, trained on large datasets. However, they have limited generalization ability to data outsid…

cs.CV2018

Hierarchical binary CNNs for landmark localization with limited resources

Adrian Bulat, Georgios Tzimiropoulos

Our goal is to design architectures that retain the groundbreaking performance of Convolutional Neural Networks (CNNs) for landmark localization and at the same time are lightweigh…

cs.LG2020

Factorized Higher-Order CNNs with an Application to Spatio-Temporal Emotion Estimation

Jean Kossaifi, Antoine Toisoul, Adrian Bulat +3

Training deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters…