6 papers
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
Shiyu Xuan, Zechao Li
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen inter…
VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification
Chao Ji, Shiyu Xuan, Zechao Li
Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric deformation. Existing methods at…
URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly Detection
Wei Luo, Peng Xing, Yunkang Cao +3
Unsupervised anomaly detection plays a pivotal role in industrial defect inspection and medical image analysis, with most methods relying on the reconstruction framework. However,…
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
Wei Tang, Xuejing Liu, Yanpeng Sun +1
The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Ex…
Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image Segmentation
Yuntian Bo, Tao Zhou, Zechao Li +2
Cross-domain few-shot medical image segmentation (CD-FSMIS) offers a promising and data-efficient solution for medical applications where annotations are severely scarce and multim…
AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic Segmentation
Jiaqi Ma, Guo-Sen Xie, Fang Zhao +1
Few-shot learning aims to recognize novel concepts by leveraging prior knowledge learned from a few samples. However, for visually intensive tasks such as few-shot semantic segment…