2 papers
cs.CV2026
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
Yuxiao Chen, Jue Wang, Zhikang Zhang +8
With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens…
cs.RO2026
MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
Xidong Li, Mingyu Guo, Chenchao Xu +5
Trajectory planning is a core task in autonomous driving, requiring the prediction of safe and comfortable paths across diverse scenarios. Integrating Multi-modal Large Language Mo…