5 papers
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
Xingming Long, Yu Liu, Zhiwei Yang +7
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or extern…
VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task
Xingming Long, Jie Zhang, Shiguang Shan +1
Most research on hallucinations in Large Vision-Language Models (LVLMs) focuses on factual description tasks that prohibit any output absent from the image. However, little attenti…
Semantic or Covariate? A Study on the Intractable Case of Out-of-Distribution Detection
Xingming Long, Jie Zhang, Shiguang Shan +1
The primary goal of out-of-distribution (OOD) detection tasks is to identify inputs with semantic shifts, i.e., if samples from novel classes are absent in the in-distribution (ID)…
Confidence Aware Learning for Reliable Face Anti-spoofing
Xingming Long, Jie Zhang, Shiguang Shan
Current Face Anti-spoofing (FAS) models tend to make overly confident predictions even when encountering unfamiliar scenarios or unknown presentation attacks, which leads to seriou…
Rethinking the Evaluation of Out-of-Distribution Detection: A Sorites Paradox
Xingming Long, Jie Zhang, Shiguang Shan +1
Most existing out-of-distribution (OOD) detection benchmarks classify samples with novel labels as the OOD data. However, some marginal OOD samples actually have close semantic con…