4 papers
Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification
Yuting Yang, Haichao Jiang, Tianming Liang +2
Referring segmentation aims to segment the target objects in images or videos based on the textual query. Despite remarkable progress over the past years, existing works always ass…
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
Quan Zhang, Zeqiang Cai, Peiming Zhao +4
Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow…
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
Weiqi Li, Quande Zhang, Ruifeng Zhai +2
Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittle…
GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization
Yirui Chen, Xudong Huang, Quan Zhang +9
The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia…