4 papers
Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration
Qi Ming, Yuyang Wang, Mingjing Zhao +7
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rel…
Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks
Qi Ming, Zheng Zhou, Haitian Yang +4
Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detection tasks, which provide a mor…
FreqTrack: Frequency Learning based Vision Transformer for RGB-Event Object Tracking
Jinlin You, Muyu Li, Xudong Zhao
Existing single-modal RGB trackers often face performance bottlenecks in complex dynamic scenes, while the introduction of event sensors offers new potential for enhancing tracking…
Event-Adaptive State Transition and Gated Fusion for RGB-Event Object Tracking
Jinlin You, Muyu Li, Xudong Zhao
Existing Vision Mamba-based RGB-Event(RGBE) tracking methods suffer from using static state transition matrices, which fail to adapt to variations in event sparsity. This rigidity…