5 papers
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati +4
Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to com…
Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
Mehdi Noroozi, Isma Hadji, Victor Escorcia +3
There has been immense progress recently in the visual quality of Stable Diffusion-based Super Resolution (SD-SR). However, deploying large diffusion models on computationally rest…
VladVA: Discriminative Fine-tuning of LVLMs
Yassine Ouali, Adrian Bulat, Alexandros Xenos +4
Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models ha…
Semi-supervised Facial Action Unit Intensity Estimation with Contrastive Learning
Enrique Sanchez, Adrian Bulat, Anestis Zaganidis +1
This paper tackles the challenging problem of estimating the intensity of Facial Action Units with few labeled images. Contrary to previous works, our method does not require to ma…
Recurrent-OctoMap: Learning State-based Map Refinement for Long-Term Semantic Mapping with 3D-Lidar Data
Li Sun, Zhi Yan, Anestis Zaganidis +2
This paper presents a novel semantic mapping approach, Recurrent-OctoMap, learned from long-term 3D Lidar data. Most existing semantic mapping approaches focus on improving semanti…