133 citations · 150 across the 8 of their papers we have counts for
10 papers
RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training
Yunshuang Nie, Bingqian Lin, Minzhe Niu +7
Pre-trained Multi-modal Large Language Models (MLLMs) provide a knowledge-rich foundation for post-training by leveraging their inherent perception and reasoning capabilities to so…
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
Jianhua Han, Meng Tian, Jiangtong Zhu +16
Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex in…
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
Bingqian Lin, Yunshuang Nie, Khun Loun Zai +10
Recent studies have revealed the potential of training open-source Large Language Models (LLMs) to unleash LLMs' reasoning ability for enhancing vision-language navigation (VLN) pe…
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
Guansong Lu, Yuanfan Guo, Jianhua Han +7
Current large-scale diffusion models represent a giant leap forward in conditional image synthesis, capable of interpreting diverse cues like text, human poses, and edges. However,…
Point2Seq: Detecting 3D Objects as Sequences
Yujing Xue, Jiageng Mao, Minzhe Niu +5
We present a simple and effective framework, named Point2Seq, for 3D object detection from point clouds. In contrast to previous methods that normally {predict attributes of 3D obj…
Voxel Transformer for 3D Object Detection
Jiageng Mao, Yujing Xue, Minzhe Niu +5
We present Voxel Transformer (VoTr), a novel and effective voxel-based Transformer backbone for 3D object detection from point clouds. Conventional 3D convolutional backbones in vo…