2 papers
cs.CV2026
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
cs.CV2025
Depth3DLane: Monocular 3D Lane Detection via Depth Prior Distillation
Dongxin Lyu, Han Huang, Cheng Tan +1
Monocular 3D lane detection is challenging due to the difficulty in capturing depth information from single-camera images. A common strategy involves transforming front-view (FV) i…