collaborators

8 papers

cs.CV2026

SAM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection

Ruichao Hou, Boyue Xu, Tongwei Ren +3

Vision foundation models have recently advanced multi-modal salient object detection (MSOD) through parameter-efficient tuning and prompt learning. However, existing Segment Anythi…

cs.CV2026

VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren +1

UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feat…

cs.CV2025

SwiTrack: Tri-State Switch for Cross-Modal Object Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren +3

Cross-modal object tracking (CMOT) is an emerging task that maintains target consistency while the video stream switches between different modalities, with only one modality availa…

cs.CV2025

Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren +3

Prompt-learning-based multi-modal trackers have made strong progress by using lightweight visual adapters to inject auxiliary-modality cues into frozen foundation models. However,…

cs.CV2025

HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection

Ruichao Hou, Xingyuan Li, Tongwei Ren +3

RGB-thermal salient object detection (RGB-T SOD) aims to identify prominent objects by integrating complementary information from RGB and thermal modalities. However, learning the…

cs.CV2025

RGB-D Tracking via Hierarchical Modality Aggregation and Distribution Network

Boyue Xu, Yi Xu, Ruichao Hou +3

The integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level featu…