activity
20242026
collaborators

6 papers

cs.CV2026

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

Nuo Chen, Lulin Liu, Zihao Li +12

Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi-agent interactions, such as…

cs.AI2026

DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents

Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7

Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…

cs.RO2026

Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis

Dayou Li, Jiuzhou Lei, Hao Wang +6

While recent foundation models have significantly advanced robotic manipulation, these systems still struggle to autonomously recover from execution errors. Current failure-learnin…

cs.CV2025

Real-Time Privacy Preservation for Robot Visual Perception

Minkyu Choi, Yunhao Yang, Neel P. Bhatt +6

Many robots (e.g., iRobot's Roomba) operate based on visual observations from live video streams, and such observations may inadvertently include privacy-sensitive objects, such as…

cs.CV2024

Towards Neuro-Symbolic Video Understanding

Minkyu Choi, Harsh Goel, Mohammad Omama +3

The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reas…

cs.CV2024

Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction

Po-han Li, Yunhao Yang, Mohammad Omama +2

Autonomous agents perceive and interpret their surroundings by integrating multimodal inputs, such as vision, audio, and LiDAR. These perceptual modalities support retrieval tasks,…