18 citations · 20 across the 3 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.CV2023★ 18 cited
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
Erfei Cui, Wenhai Wang, Zhiqi Li +7
Large language models (LLMs) have opened up new possibilities for intelligent agents, endowing them with human-like thinking and cognitive abilities. In this work, we delve into th…
cs.CV2023★ 2 cited
InstructSeq: Unifying Vision Tasks with Instruction-conditioned Multi-modal Sequence Generation
Rongyao Fang, Shilin Yan, Zhaoyang Huang +4
Empowering models to dynamically accomplish tasks specified through natural language instructions represents a promising path toward more capable and general artificial intelligenc…