5 papers
Language-based Trial and Error Falls Behind in the Era of Experience
Haoyu Wang, Guozheng Ma, Shugang Cui +7
While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic or spatial tasks) remains limite…
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
Ziyang Gong, Zehang Luo, Anke Tang +21
Universal embodied intelligence demands robust generalization across heterogeneous embodiments, such as autonomous driving, robotics, and unmanned aerial vehicles (UAVs). However,…
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
Erfei Cui, Wenhai Wang, Zhiqi Li +7
Large language models (LLMs) have opened up new possibilities for intelligent agents, endowing them with human-like thinking and cognitive abilities. In this work, we delve into th…
Demystify Transformers & Convolutions in Modern Image Deep Networks
Xiaowei Hu, Min Shi, Weiyun Wang +9
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these adva…
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Hao Li, Changyao Tian, Jie Shao +8
The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent effort…