3 papers
cs.CV2025
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
Mingliang Zhai, Cheng Li, Zengyuan Guo +7
The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. Howeve…
cs.CV2024
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
Hannan Lu, Xiaohe Wu, Shudong Wang +5
Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Exist…
cs.CV2024
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
Yukun Zhai, Xiaoqiang Zhang, Xiameng Qin +3
End-to-end text spotting is a vital computer vision task that aims to integrate scene text detection and recognition into a unified framework. Typical methods heavily rely on Regio…