2 papers
cs.CV2025
MTNet: Learning modality-aware representation with transformer for RGBT tracking
Ruichao Hou, Boyue Xu, Tongwei Ren +1
The ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tr…
cs.CV2025
Spatial-Temporal Human-Object Interaction Detection
Xu Sun, Yunqing He, Tongwei Ren +1
In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (H…