6 papers
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
Qinghongbing Xie, Zhaoyuan Xia, Feng Zhu +4
Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Exist…
AssemMate: Graph-Based LLM for Robotic Assembly Assistance
Qi Zheng, Chaoran Zhang, Zijian Liang +5
Large Language Model (LLM)-based robotic assembly assistance has gained significant research attention. It requires the injection of domain-specific knowledge to guide the assembly…
Embodied intelligent industrial robotics: Framework and techniques
Chaoran Zhang, Chenhao Zhang, Zhaobo Xu +4
The combination of embodied intelligence and robots has great prospects and is becoming increasingly common. In order to work more efficiently, accurately, reliably, and safely in…
Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
Xiaoming Zhu, Xu Huang, Qinghongbing Xie +8
Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, w…
DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding
Qinghongbing Xie, Zijian Liang, Fuhao Li +1
Effective scene representation is critical for the visual grounding ability of representations, yet existing methods for 3D Visual Grounding are often constrained. They either only…
Demonstrating DVS: Dynamic Virtual-Real Simulation Platform for Mobile Robotic Tasks
Zijie Zheng, Zeshun Li, Yunpeng Wang +2
With the development of embodied artificial intelligence, robotic research has increasingly focused on complex tasks. Existing simulation platforms, however, are often limited to i…