6 papers
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
Nuo Chen, Lulin Liu, Zihao Li +12
Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi-agent interactions, such as…
DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents
Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7
Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…
Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis
Dayou Li, Jiuzhou Lei, Hao Wang +6
While recent foundation models have significantly advanced robotic manipulation, these systems still struggle to autonomously recover from execution errors. Current failure-learnin…
Real-Time Privacy Preservation for Robot Visual Perception
Minkyu Choi, Yunhao Yang, Neel P. Bhatt +6
Many robots (e.g., iRobot's Roomba) operate based on visual observations from live video streams, and such observations may inadvertently include privacy-sensitive objects, such as…
Towards Neuro-Symbolic Video Understanding
Minkyu Choi, Harsh Goel, Mohammad Omama +3
The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reas…
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
Po-han Li, Yunhao Yang, Mohammad Omama +2
Autonomous agents perceive and interpret their surroundings by integrating multimodal inputs, such as vision, audio, and LiDAR. These perceptual modalities support retrieval tasks,…