5 papers
QuantWAMs: Calibrating at the Right Granularity for World Action Models
Jiacheng Zhou, Jinfan Lv, Ruixuan Li +4
World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-…
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
Zhaoyang Yang, Yurun Jin, Lizhe Qi +2
Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models joi…
EmoScene: A Dual-space Dataset for Controllable Affective Image Generation
Li He, Longtai Zhang, Wenqiang Zhang +2
Text-to-image diffusion models achieve high visual fidelity, yet fine-grained affective control remains difficult because textual emotion cues often fail to specify the visual perc…
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
Rushuai Yang, Hecheng Wang, Zhichao Wu +11
We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is l…
MMARD: Improving the Min-Max Optimization Process in Adversarial Robustness Distillation
Yuzheng Wang, Zhaoyu Chen, Dingkang Yang +2
Adversarial Robustness Distillation (ARD) is a promising task to boost the robustness of small-capacity models with the guidance of the pre-trained robust teacher. The ARD can be s…