9 papers · 1 filter
Can We Perform Online RL for Image Editing without Editing Rewards?
Qichao Ma, Jikang Cheng, Ling Liang +3
Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly triplet supervision…
Brain-Inspired Multimodal Spiking Neural Network for Image-Text Retrieval
Xintao Zong, Xian Zhong, Wenxuan Liu +3
Spiking neural networks (SNNs) have recently shown strong potential in unimodal visual and textual tasks, yet building a directly trained, low-energy, and high-performance SNN for…
Driving in Spikes: An Entropy-Guided Object Detector for Spike Cameras
Ziyan Liu, Qi Su, Lulu Tang +2
Object detection in autonomous driving suffers from motion blur and saturation under fast motion and extreme lighting. Spike cameras, offer microsecond latency and ultra high dynam…
HAD: Hierarchical Asymmetric Distillation to Bridge Spatio-Temporal Gaps in Event-Based Object Tracking
Yao Deng, Xian Zhong, Wenxuan Liu +3
RGB cameras excel at capturing rich texture details with high spatial resolution, whereas event cameras offer exceptional temporal resolution and a high dynamic range (HDR). Levera…
SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams
Zhuoheng Gao, Yihao Li, Jiyao Zhang +7
Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microse…
SPKLIP: Aligning Spike Video Streams with Natural Language
Yongchang Gao, Meiling Jin, Zhaofei Yu +2
Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) w…