6 papers
Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning
Xinchuan Qiu, Yi Yu
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generat…
Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis
Yi Yu, Xinchuan Qiu
Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive…
SHERPA: Seam-aware Harmonized ERP Adaptation for Open-Domain 360 Panorama Generation
Jungwoon Kang, Jaehun Kim, Yiwon Yu +3
Panoramic imagery is increasingly used in world-generation, games, and simulation, where users may need not only photorealistic scenes but also stylized and non-photorealistic envi…
Universal Adversarial Attacks against Closed-Source MLLMs via Target-View Routed Meta Optimization
Hui Lu, Yi Yu, Yiming Yang +6
Targeted adversarial attacks on closed-source multimodal large language models (MLLMs) have been increasingly explored under black-box transfer, yet prior methods are predominantly…
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
Hui Lu, Yi Yu, Yiming Yang +5
Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single…
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
Hui Lu, Yi Yu, Shijian Lu +4
Temporal Action Detection (TAD) aims to identify and localize actions by determining their starting and ending frames within untrimmed videos. Recent Structured State-Space Models…