Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
SAM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection
Ruichao Hou, Boyue Xu, Tongwei Ren +3
Vision foundation models have recently advanced multi-modal salient object detection (MSOD) through parameter-efficient tuning and prompt learning. However, existing Segment Anythi…
cs.CV2025
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
Boyue Xu, Ruichao Hou, Tongwei Ren +3
Cross-modal object tracking (CMOT) is an emerging task that maintains target consistency while the video stream switches between different modalities, with only one modality availa…
cs.CV2025
Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking
Boyue Xu, Ruichao Hou, Tongwei Ren +3
Prompt-learning-based multi-modal trackers have made strong progress by using lightweight visual adapters to inject auxiliary-modality cues into frozen foundation models. However,…