4 papers
Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
Yimei Zhang, Guojiang Shen, Kaili Ning +4
Region representation learning plays a pivotal role in urban computing by extracting meaningful features from unlabeled urban data. Analogous to how perceived facial age reflects a…
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
Boyue Xu, Ruichao Hou, Tongwei Ren +3
Cross-modal object tracking (CMOT) is an emerging task that maintains target consistency while the video stream switches between different modalities, with only one modality availa…
HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection
Ruichao Hou, Xingyuan Li, Tongwei Ren +3
RGB-thermal salient object detection (RGB-T SOD) aims to identify prominent objects by integrating complementary information from RGB and thermal modalities. However, learning the…
Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking
Boyue Xu, Ruichao Hou, Tongwei Ren +3
Prompt-learning-based multi-modal trackers have made strong progress by using lightweight visual adapters to inject auxiliary-modality cues into frozen foundation models. However,…