collaborators

5 papers

cs.CV2026

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Kailing Li, Qi'ao Xu, Tianwen Qian +3

Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and spatiotemporal reasoning in dynamic e…

cs.CV2026

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos

Qi'ao Xu, Tianwen Qian, Yuqian Fu +5

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG…

cs.CV2026

EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering

Yanjun Li, Yuqian Fu, Tianwen Qian +5

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmar…

cs.CV2025

HSACNet: Hierarchical Scale-Aware Consistency Regularized Semi-Supervised Change Detection

Qi'ao Xu, Pengfei Wang, Yanjun Li +2

Semi-supervised change detection (SSCD) aims to detect changes between bi-temporal remote sensing images by utilizing limited labeled data and abundant unlabeled data. Existing met…

cs.CV2025

TS-PCL: Plug-and-Play Dual Contrastive Learning for Vision-Guided Medical Time Series Classification

Qi'ao Xu, Pengfei Wang, Bo Zhong +4

Medical time series (MedTS) classification is pivotal for intelligent healthcare, yet its efficacy is severely limited by poor cross-subject generation due to the profound cross-in…