From the 1 of 7 linked papers with an AI index.
7 papers
SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
Haidong Cao, Wenjun Cao, Quanhao Li +5
The paper introduces SegDiff, a closed-loop visuomotor policy that segments demonstrations into motion segments and uses diffusion models to predict continuous trajectories to the…
Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
Xingyao Lin, Xinghao Zhu, Tianyi Lu +6
Embodied agents are intelligent systems designed to perceive, reason, and act within the physical world. While the robotics community has long strived to build such versatile agent…
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Xintong Hu, Xuhong Huang, Jinyu Zhang +11
Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However…
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Qiuyue Wang, Mingsheng Li, Jian Guan +37
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generali…
Unify Robot Actions in Camera Frame
Sicheng Xie, Lingchen Meng, Zijie Diao +9
Cross-embodiment robot learning requires a unified action representation with consistent semantics across robot platforms. Existing representations suffer from platform-specific in…
Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
Jiaqi Leng, Shuyuan Tu, Haidong Cao +4
Human preference alignment presents a critical yet underexplored challenge for diffusion models in text-to-3D generation. Existing solutions typically require task-specific fine-tu…