activity
20232026
most citedVladVA: Discriminative Fine-tuning of LVLMs

2 citations · 5 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2026

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati +4

Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to com…

cs.CV2026

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas +2

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates…

cs.CV2026

More Images, More Problems? A Controlled Analysis of VLM Failure Modes

Anurag Das, Adrian Bulat, Alberto Baldrati +4

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities, yet their proficiency in understanding and reasoning over multiple images remains largely unexplored…

cs.CV2024★ 2 cited

VladVA: Discriminative Fine-tuning of LVLMs

Yassine Ouali, Adrian Bulat, Alexandros Xenos +4

Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models ha…

cs.CV2024

Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing

Ioannis Maniadis Metaxas, Georgios Tzimiropoulos, Ioannis Patras

Self-supervised learning has recently emerged as the preeminent pretraining paradigm across and between modalities, with remarkable results. In the image domain specifically, group…

cs.CV2023★ 1 cited

Aligned Unsupervised Pretraining of Object Detectors with Self-training

Ioannis Maniadis Metaxas, Adrian Bulat, Ioannis Patras +2

The unsupervised pretraining of object detectors has recently become a key component of object detector training, as it leads to improved performance and faster convergence during…