12 papers
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Xingcheng Zhou, Hao Guo, Rui Song +5
Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses unde…
SegRGB-X: General RGB-X Semantic Segmentation Model
Jiong Liu, Yingjie Xu, Xingcheng Zhou +4
Semantic segmentation across arbitrary sensor modalities faces significant challenges due to diverse sensor characteristics, and the traditional configurations for this task result…
Energy-Aware Imitation Learning for Steering Prediction Using Events and Frames
Hu Cao, Jiong Liu, Xingzhuo Yan +5
In autonomous driving, relying solely on frame-based cameras can lead to inaccuracies caused by factors like long exposure times, high-speed motion, and challenging lighting condit…
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
Dingyi Zhou, Mu He, Zhuowei Fang +4
We introduce AffordanceGrasp-R1, a reasoning-driven affordance segmentation framework for robotic grasping that combines a chain-of-thought (CoT) cold-start strategy with reinforce…
Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
Zebin Jiang, Tianle Jin, Xiangtong Yao +2
Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent resea…
URNet: Uncertainty-aware Refinement Network for Event-based Stereo Depth Estimation
Yifeng Cheng, Alois Knoll, Hu Cao
Event cameras provide high temporal resolution, high dynamic range, and low latency, offering significant advantages over conventional frame-based cameras. In this work, we introdu…