activity
20242026
collaborators

8 papers

cs.CV2026

Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

Jianzhe Gao, Churan Wang, Weiyi Zhang +5

Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on an end-to-end learning paradi…

cs.RO2026

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking

Kui Wu, Hao Chen, Jinzhu Han +6

Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platfor…

cs.AI2025

UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

Fangwei Zhong, Kui Wu, Churan Wang +4

We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environm…

cs.CV2025

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

Kui Wu, Shuhang Xu, Hao Chen +4

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tr…

cs.CV2025

Hierarchical Instruction-aware Embodied Visual Tracking

Kui Wu, Hao Chen, Churan Wang +4

User-Centric Embodied Visual Tracking (UC-EVT) presents a novel challenge for reinforcement learning-based models due to the substantial gap between high-level user instructions an…

cs.CV2025

Autoregressive Sequence Modeling for 3D Medical Image Representation

Siwen Wang, Churan Wang, Fei Gao +4

Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse…