13 papers
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
Jian Zhu, Jianjun Zhang, Taiyi Su +10
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learni…
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
Taiyi Su, Jian Zhu, Tianjian Wang +9
Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and househ…
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
Shengjun Zhang, Zhang Zhang, Simin Huang +11
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between…
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
Hanyang Wang, Yiyang Liu, Jiawei Chi +3
Classifier-Free Guidance (CFG) has emerged as a central approach for enhancing semantic alignment in flow-based diffusion models. In this paper, we explore a unified framework call…
Unsupervised Multi-Attention Meta Transformer for Rotating Machinery Fault Diagnosis
Hanyang Wang, Yuxuan Yang, Hongjun Wang +1
The intelligent fault diagnosis of rotating mechanical equipment usually requires a large amount of labeled sample data. However, in practical industrial applications, acquiring en…
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
Fangfu Liu, Hao Li, Jiawei Chi +4
Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene…