activity
20222026
most citedSelf-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks

15 citations · 27 across the 8 of their papers we have counts for

collaborators

8 papers

cs.AI2026

iFLYTEK-Embodied-Omni Technical Report

Yuan Zhang, Jingfei Ni, Guanchen Lu +12

General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. E…

cs.RO2026

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Yuan Zhang, Shiqi Zhang, Yedong Shen +11

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…

cs.SE2024

ROCAS: Root Cause Analysis of Autonomous Driving Accidents via Cyber-Physical Co-mutation

Shiwei Feng, Yapeng Ye, Qingkai Shi +5

As Autonomous driving systems (ADS) have transformed our daily life, safety of ADS is of growing significance. While various testing approaches have emerged to enhance the ADS reli…

cs.RO2024

EffiTune: Diagnosing and Mitigating Training Inefficiency for Parameter Tuner in Robot Navigation System

Shiwei Feng, Xuan Chen, Zikang Xiong +5

Robot navigation systems are critical for various real-world applications such as delivery services, hospital logistics, and warehouse management. Although classical navigation met…

cs.CV2024★ 15 cited

Self-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks

Zhiyuan Cheng, Cheng Han, James Liang +3

Monocular Depth Estimation (MDE) plays a vital role in applications such as autonomous driving. However, various attacks target MDE models, with physical attacks posing significant…

cs.CV2023★ 3 cited

Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection

Zhiyuan Cheng, Hongjun Choi, James Liang +5

Multi-sensor fusion (MSF) is widely used in autonomous vehicles (AVs) for perception, particularly for 3D object detection with camera and LiDAR sensors. The purpose of fusion is t…