16 papers
Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
Lingrui Li, Nan Pu, Dong Zhao +4
Test-time adaptation (TTA) aims to mitigate distribution shifts by adapting models with unlabeled target data at inference time. While TTA with vision-language models (VLMs) has sh…
Technical Report for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Pretraining-Diverse Ensemble of Foundation Vision Encoders for Robust Outdoor Scene Understanding
Boyan Wang, Yongxi Huang, Wenjing Li +4
This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor scenes from four camera platf…
CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection
Jinjie Shen, Yaxiong Wang, Yujiao Wu +5
The rapid rise of generative AI has made multimodal fake news increasingly realistic and pervasive, posing severe threats to public trust and social stability. Existing detection m…
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
Jinjie Shen, Zheng Huang, Yuchen Zhang +7
Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-con…
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
Jinjie Shen, Jing Wu, Yaxiong Wang +5
Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinform…
The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery
Haiyang Zheng, Nan Pu, Yaqi Cai +4
Generalized Category Discovery (GCD) leverages labeled data to categorize unlabeled samples from known or unknown classes. Most previous methods jointly optimize supervised and uns…