From the 1 of 12 linked papers with an AI index.
12 papers
Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation
Donglin Yang, Haoran Chen, Xingyu Chen +6
Mobile manipulation is a key capability for embodied intelligence, enabling robots to accomplish complex multi-stage tasks in open-world environments. However, mobile manipulation…
CR-Solver: GPU-Accelerated Kinematics Solver for Tendon-driven Continuum Robots
Heqing Yang, Yang Yi, Linqing Zhong +2
The paper introduces CR-Solver, a GPU‑accelerated optimization framework that solves inverse kinematics, path following, and trajectory planning for tendon‑driven continuum robots…
EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots
Heqing Yang, Yang Yi, Liyao Wang +8
We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and…
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Hohin Kwan, Hongyu Li, Ray Zhang +5
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
Le Zhuo, Songhao Han, Yuandong Pu +8
While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and…
PICABench: How Far Are We from Physically Realistic Image Editing?
Yuandong Pu, Le Zhuo, Songhao Han +10
Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond complet…