collaborators

6 papers

cs.RO2026

VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments

Haolin Yang, Yuxing Long, Zihan Yang +1

Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset constr…

cs.CV2026

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

Jiyao Zhang, Mingxu Zhang, Yitong Peng +8

Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric bench…

cs.RO2026

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

Haolin Yang, Yuxing Long, Zhuoyuan Yu +8

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigati…

cs.RO2025

RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals

Yuzheng Gao, Yuxing Long, Lei Kang +8

Existing appliance assets suffer from poor rendering, incomplete mechanisms, and misalignment with manuals, leading to simulation-reality gaps that hinder appliance manipulation de…

cs.RO2025

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

Zhuoyuan Yu, Yuxing Long, Zihan Yang +4

Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capabili…

cs.CV2025

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation

Yuxing Long, Jiyao Zhang, Mingjie Pan +3

Correct use of electrical appliances has significantly improved human life quality. Unlike simple tools that can be manipulated with common sense, different parts of electrical app…