2 papers
cs.CV2026
ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation
Yuanjia Li, Tianyang Xu, Tao Zhou +3
Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large languag…
cs.CV2026
Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space
Huan Kang, Hui Li, Tianyang Xu +3
Infrared and visible image fusion aims to integrate complementary modalities, while existing Euclidean methods impose rigid distance metrics that distort multi-modal interactions a…