10 papers
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning
Yanzhe Tang, Xinyu Shao, Yuxuan Hu +6
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This…
ATG-MoE: Autoregressive trajectory generation with mixture-of-experts for assembly skill learning
Weihang Huang, Chaoran Zhang, Xiaoxin Deng +4
Flexible manufacturing requires robot systems that can adapt to constantly changing tasks, objects, and environments. However, traditional robot programming is labor-intensive and…
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
Qinghongbing Xie, Zhaoyuan Xia, Feng Zhu +4
Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Exist…
AssemMate: Graph-Based LLM for Robotic Assembly Assistance
Qi Zheng, Chaoran Zhang, Zijian Liang +5
Large Language Model (LLM)-based robotic assembly assistance has gained significant research attention. It requires the injection of domain-specific knowledge to guide the assembly…
SculptDrug : A Spatial Condition-Aware Bayesian Flow Model for Structure-based Drug Design
Qingsong Zhong, Haomin Yu, Yan Lin +3
Structure-Based drug design (SBDD) has emerged as a popular approach in drug discovery, leveraging three-dimensional protein structures to generate drug ligands. However, existing…
Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
Xiaoming Zhu, Xu Huang, Qinghongbing Xie +8
Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, w…