collaborators

6 papers

cs.CV2026

Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision

Beibei Zhang, Chao Xu, Jun Lan +4

Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise i…

cs.CV2025

MTNet: Learning modality-aware representation with transformer for RGBT tracking

Ruichao Hou, Boyue Xu, Tongwei Ren +1

The ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tr…

cs.CV2025

Spatial-Temporal Human-Object Interaction Detection

Xu Sun, Yunqing He, Tongwei Ren +1

In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (H…

cs.CV2025

RGB-D Tracking via Hierarchical Modality Aggregation and Distribution Network

Boyue Xu, Yi Xu, Ruichao Hou +3

The integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level featu…

cs.CV2025

RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory

Boyue Xu, Ruichao Hou, Tongwei Ren +1

The RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues of depth modality, boosting the…

cs.MM2025

KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection

Xingyuan Li, Ruichao Hou, Tongwei Ren +1

Existing RGB-thermal salient object detection (RGB-T SOD) methods aim to identify visually significant objects by leveraging both RGB and thermal modalities to enable robust perfor…