activity
20242026
collaborators

7 papers

cs.CV2026

FaceXBench: Evaluating Multimodal LLMs on Face Understanding

Kartik Narayan, Vibashan VS, Vishal M. Patel

Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding h…

cs.LG2025

PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models

Jenny Schmalfuss, Nadine Chang, Vibashan VS +3

Vision language models (VLMs) respond to user-crafted text prompts and visual inputs, and are applied to numerous real-world problems. VLMs integrate visual modalities with large l…

cs.CV2025

Certainty and Uncertainty Guided Active Domain Adaptation

Bardia Safaei, Vibashan VS, Vishal M. Patel

Active Domain Adaptation (ADA) adapts models to target domains by selectively labeling a few target samples. Existing ADA methods prioritize uncertain samples but overlook confiden…

cs.CV2025

FaceXFormer: A Unified Transformer for Facial Analysis

Kartik Narayan, Vibashan VS, Rama Chellappa +1

In this work, we introduce FaceXFormer, an end-to-end unified transformer model capable of performing ten facial analysis tasks within a single framework. These tasks include face…

cs.CV2025

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Zhiqi Li, Guo Chen, Shilong Liu +24

Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most…

cs.CV2025

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models

Yasiru Ranasinghe, Vibashan VS, James Uplinger +2

Automatic target recognition (ATR) plays a critical role in tasks such as navigation and surveillance, where safety and accuracy are paramount. In extreme use cases, such as milita…