works on

From the 1 of 29 linked papers with an AI index.

activity
20242026
collaborators

29 papers

cs.RO2026

Towards Human-level Dexterous Teleoperation

Puhao Li, Zeyuan Chen, Yingying Wu +9

The paper presents TeleDexter, a hand‑object co‑tracking controller that learns to map human teleoperation intent into low‑level contact actions for dexterous robot hands, achievin…

cs.CV2026

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Mengya Liu, Baoxiong Jia, Jiangyong Huang +2

Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-qualit…

cs.CV2026

3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding

Xiongkun Linghu, Jiangyong Huang, Baoxiong Jia +1

Reinforcement Learning with Verifiable Rewards ( RLVR ) has emerged as a transformative paradigm for enhancing the reasoning capabilities of Large Language Models ( LLMs), yet its…

cs.CV2026

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

Yixin Chen, Yaowei Zhang, Huangyue Yu +9

Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully desi…

cs.CV2026

LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning

Jiangyong Huang, Xiaojian Ma, Xiongkun Linghu +6

Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reas…

cs.RO2026

OmniClone: Engineering a Robust, All-Rounder Whole-Body Humanoid Teleoperation System

Yixuan Li, Le Ma, Yutang Lin +8

Whole-body humanoid teleoperation enables humans to remotely control humanoid robots, serving as both a real-time operational tool and a scalable engine for collecting demonstratio…