Publications (53)
Subpixel Heatmap Regression for Facial Landmark Localization
Adrian Bulat, Enrique Sanchez, Georgios Tzimiropoulos
Deep Learning models based on heatmap regression have revolutionized the task of facial landmark localization with existing models working robustly under large poses, non-uniform i…
REST: REtrieve & Self-Train for generative action recognition
Adrian Bulat, Enrique Sanchez, Brais Martinez +1
This work is on training a generative action/video recognition model whose output is a free-form action-specific caption describing the video (rather than an action class label). A…
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas +2
Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates…
Tensor Dropout for Robust Learning
Arinbjörn Kolbeinsson, Jean Kossaifi, Yannis Panagakis +4
CNNs achieve remarkable performance by leveraging deep, over-parametrized architectures, trained on large datasets. However, they have limited generalization ability to data outsid…
Hierarchical binary CNNs for landmark localization with limited resources
Adrian Bulat, Georgios Tzimiropoulos
Our goal is to design architectures that retain the groundbreaking performance of Convolutional Neural Networks (CNNs) for landmark localization and at the same time are lightweigh…
Factorized Higher-Order CNNs with an Application to Spatio-Temporal Emotion Estimation
Jean Kossaifi, Antoine Toisoul, Adrian Bulat +3
Training deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters…