6 papers · 1 filter
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
Kartik Narayan, Vibashan VS, Vishal M. Patel
Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding h…
Certainty and Uncertainty Guided Active Domain Adaptation
Bardia Safaei, Vibashan VS, Vishal M. Patel
Active Domain Adaptation (ADA) adapts models to target domains by selectively labeling a few target samples. Existing ADA methods prioritize uncertain samples but overlook confiden…
FaceXFormer: A Unified Transformer for Facial Analysis
Kartik Narayan, Vibashan VS, Rama Chellappa +1
In this work, we introduce FaceXFormer, an end-to-end unified transformer model capable of performing ten facial analysis tasks within a single framework. These tasks include face…
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen, Shilong Liu +24
Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most…
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
Yasiru Ranasinghe, Vibashan VS, James Uplinger +2
Automatic target recognition (ATR) plays a critical role in tasks such as navigation and surveillance, where safety and accuracy are paramount. In extreme use cases, such as milita…
SegFace: Face Segmentation of Long-Tail Classes
Kartik Narayan, Vibashan VS, Vishal M. Patel
Face parsing refers to the semantic segmentation of human faces into key facial regions such as eyes, nose, hair, etc. It serves as a prerequisite for various advanced applications…