collaborators

6 papers

cs.CV2026

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

Yuanjia Li, Tianyang Xu, Tao Zhou +3

Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large languag…

cs.CV2026

EvaNet: Towards More Efficient and Consistent Infrared and Visible Image Fusion Assessment

Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu +4

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, ofte…

cs.CV2026

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

Yuanbo Li, Tianyang Xu, Cong Hu +3

The rapid progress of Multi-Modal Large Language Models (MLLMs) has significantly advanced downstream applications. However, this progress also exposes serious transferable adversa…

cs.CV2026

Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

Yuanbo Li, Tianyang Xu, Cong Hu +3

With the rapid advancement and widespread application of vision-language pre-training (VLP) models, their vulnerability to adversarial attacks has become a critical concern. In gen…

cs.CV2026

Omni Survey for Multimodality Analysis in Visual Object Tracking

Zhangyong Tang, Tianyang Xu, Xuefeng Zhu +6

The development of smart cities has led to the generation of massive amounts of multi-modal data in the context of a range of tasks that enable a comprehensive monitoring of the sm…

cs.CV2025

Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and Benchmarking

Zhangyong Tang, Tianyang Xu, Xuefeng Zhu +4

Unifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking sy…