collaborators

6 papers

cs.RO2026

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

Xinchuan Qiu, Yi Yu

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generat…

cs.RO2026

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Yi Yu, Xinchuan Qiu

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive…

cs.CV2026

SHERPA: Seam-aware Harmonized ERP Adaptation for Open-Domain 360 Panorama Generation

Jungwoon Kang, Jaehun Kim, Yiwon Yu +3

Panoramic imagery is increasingly used in world-generation, games, and simulation, where users may need not only photorealistic scenes but also stylized and non-photorealistic envi…

cs.AI2026

Universal Adversarial Attacks against Closed-Source MLLMs via Target-View Routed Meta Optimization

Hui Lu, Yi Yu, Yiming Yang +6

Targeted adversarial attacks on closed-source multimodal large language models (MLLMs) have been increasingly explored under black-box transfer, yet prior methods are predominantly…

cs.CV2026

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

Hui Lu, Yi Yu, Yiming Yang +5

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single…

cs.CV2026

MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection

Hui Lu, Yi Yu, Shijian Lu +4

Temporal Action Detection (TAD) aims to identify and localize actions by determining their starting and ending frames within untrimmed videos. Recent Structured State-Space Models…