works on

From the 1 of 15 linked papers with an AI index.

activity
20242026
collaborators

15 papers

cs.AI2026

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

Kun Yu, Jianhua Yang, Yixiang Chen +7

The paper introduces UESF-Bench, a large-scale benchmark for unified embodied seeking and following of humans, and presents SeekFollow-VLA, a vision‑language‑action framework that…

cs.RO2026

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

Yuan Xu, Yixiang Chen, Kai Wang +5

Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…

cs.RO2026

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Peiyan Li, Yixiang Chen, Yuan Xu +13

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. The…

cs.RO2026

EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation

Yuan Xu, Jiabing Yang, Xiaofeng Wang +16

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-…

cs.CV2026

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

Tao Yu, Haopeng Jin, Hao Wang +18

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Ope…

cs.CV2026

Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization

Tao Yu, Yujia Yang, Haopeng Jin +17

Traditional video retrieval benchmarks focus on matching precise descriptions to closed video pools, failing to reflect real-world searches characterized by fuzzy, multi-dimensiona…