4 papers
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
Shiyu Xuan, Zechao Li
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen inter…
VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification
Chao Ji, Shiyu Xuan, Zechao Li
Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric deformation. Existing methods at…
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
Shiyu Xuan, Dongkai Wang, Zechao Li +1
Zero-shot Human-object interaction (HOI) detection aims to locate humans and objects in images and recognize their interactions. While advances in open-vocabulary object detection…
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
Shiyu Xuan, Zechao Li, Jinhui Tang
Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing g…