5 papers
Towards the Harness of Embodied Agents
Qi Wang, Tianyi Wang, Chengyang Li +6
The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether t…
Semantic Anchoring for Robotic Action Representations
Yuan Xu, Youheng Shi, Chengyang Li +2
The paper studies how fine‑tuning vision‑language‑action models for robots can degrade the semantic structure of their action representations, and proposes a plug‑and‑play anchorin…
GazeVLA: Learning Human Intention for Robotic Manipulation
Chengyang Li, Kaiyi Xiong, Yuan Xu +3
Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works…
GenDet: Painting Colored Bounding Boxes on Images via Diffusion Model for Object Detection
Chen Min, Chengyang Li, Fanjie Kong +3
This paper presents GenDet, a novel framework that redefines object detection as an image generation task. In contrast to traditional approaches, GenDet adopts a pioneering approac…
C-DGPA: Class-Centric Dual-Alignment Generative Prompt Adaptation
Chao Li, Dasha Hu, Chengyang Li +2
Unsupervised Domain Adaptation transfers knowledge from a labeled source domain to an unlabeled target domain. Directly deploying Vision-Language Models (VLMs) with prompt tuning i…