12 papers
SpikingMOT: A Spike-Driven Multi-Object Tracker
Yiding Sun, Xiangyang Yang, Dongxu Zhang +7
Multi-object tracking (MOT) plays a fundamental role in visual perception, where accurate trajectory prediction is essential for reliable target association under complex motion pa…
SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams
Zhuoheng Gao, Yihao Li, Jiyao Zhang +7
Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microse…
SPKLIP: Aligning Spike Video Streams with Natural Language
Yongchang Gao, Meiling Jin, Zhaofei Yu +2
Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) w…
SpikeGrasp: A Benchmark for 6-DoF Grasp Pose Detection from Stereo Spike Streams
Zhuoheng Gao, Jiyao Zhang, Zhiyong Xie +5
Most robotic grasping systems rely on converting sensor data into explicit 3D point clouds, which is a computational step not found in biological intelligence. This paper explores…
: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
Kang Chen, Zhihao Liu, Tonghe Zhang +11
Vision-Language-Action (VLA) models enable robots to understand and perform complex tasks from multimodal input. Although recent work explores using reinforcement learning (RL) to…
Inner-Probe: Discovering Copyright-related Data Generation in LLM Architecture
Qichao Ma, Rui-Jie Zhu, Peiye Liu +8
Large Language Models (LLMs) utilize extensive knowledge databases and show powerful text generation ability. However, their reliance on high-quality copyrighted datasets raises co…