collaborators

20 papers

cs.CV2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Zongchuang Zhao, Xin Zhou, Tianyang Xu +5

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test-time future imag…

cs.CV2026

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

Yuanjia Li, Tianyang Xu, Tao Zhou +3

Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large languag…

cs.CV2026

Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach

Yingjie Dai, Tianyang Xu, Yanglin Deng +2

Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume co…

cs.CV2026

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

Huan Kang, Hui Li, Tianyang Xu +3

Infrared and visible image fusion aims to integrate complementary modalities, while existing Euclidean methods impose rigid distance metrics that distort multi-modal interactions a…

cs.CV2026

EvaNet: Towards More Efficient and Consistent Infrared and Visible Image Fusion Assessment

Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu +4

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, ofte…

cs.CV2026

Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion

Yanglin Deng, Tianyang Xu, Chunyang Cheng +3

Infrared and visible image fusion(IVIF) combines complementary modalities while preserving natural textures and salient thermal signatures. Existing solutions predominantly rely on…