7 papers
JWST MIRI Medium Resolution Spectrometer Point Fixed Pattern Corrections: Cleaner and Higher Signal-to-Noise Spectra of Point Sources
Karl D. Gordon, David R. Law
The JWST Mid-Infrared Instrument Medium Resolution Spectrometer provides the capability to obtain spectra from 5-28 micron. The JWST data reduction pipeline removes the majority bu…
AVA: Attentive VLM Agent for Mastering StarCraft II
Weiyu Ma, Yuqian Fu, Zecheng Zhang +2
We introduce AVACraft, a multimodal StarCraft II benchmark supporting both Multi-Agent Reinforcement Learning (MARL) and Vision-Language Model (VLM) paradigms. Unlike SMAC-family e…
X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations
Maximus A. Pace, Prithwish Dan, Chuanruo Ning +5
Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for dire…
Implicit State Estimation via Video Replanning
Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu +5
Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These re…
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real
Prithwish Dan, Kushal Kedia, Angela Chao +4
Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment appro…
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
Chuanruo Ning, Kuan Fang, Wei-Chiu Ma
Recent advancements in open-world robot manipulation have been largely driven by vision-language models (VLMs). While these models exhibit strong generalization ability in high-lev…