activity
20232026
collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects

Tianjing Hao, Haiyu Lan, Angsong Li +6

Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-language detectors recover long-tail categories and small, fine-grain…

cs.RO2026

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

Jiarui Yang, Yehao Lu, Yuning Su +9

Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent phys…

cs.RO2026

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

Jiarui Yang, Bin Zhu, Jingjing Chen +4

Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for ad…

cs.RO2026

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

Jiarui Yang, Wen Huang, Jiale Zhang +2

Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expe…

cs.RO2026

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

Jiarui Yang, Yehao Lu, Yuning Su +9

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that joint…