collaborators

6 papers

cs.RO2026

VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments

Haolin Yang, Yuxing Long, Zihan Yang +1

Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset constr…

cs.CL2026

Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents

Xushuo Tang, Junhe Zhang, Zihan Yang +4

Recent LLM role-playing systems build character agents from novels by extracting characters, scenes, and relations. Yet long-narrative role-playing suffers from two failures: Factu…

cs.RO2026

Playful Agentic Robot Learning

Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reu…

cs.CV2026

Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

Shuyuan Tu, Qi Tian, Zihan Yang +9

Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cau…

cs.RO2026

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

Haolin Yang, Yuxing Long, Zhuoyuan Yu +8

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigati…

cs.RO2025

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

Zhuoyuan Yu, Yuxing Long, Zihan Yang +4

Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capabili…