5 papers
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Yuran Wang, Siqiao Huang, Mingleyang Li +21
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are…
WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models
Chenhao Zhang, Hanyu Zhao, Hang Cheng +2
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world e…
Instance camera focus prediction for crystal agglomeration classification
Xiaoyu Ji, Chenhao Zhang, Tyler James Downard +3
Agglomeration refers to the process of crystal clustering due to interparticle forces. Crystal agglomeration analysis from microscopic images is challenging due to the inherent lim…
ICP-3DGS: SfM-free 3D Gaussian Splatting for Large-scale Unbounded Scenes
Chenhao Zhang, Yezhi Shen, Fengqing Zhu
In recent years, neural rendering methods such as NeRFs and 3D Gaussian Splatting (3DGS) have made significant progress in scene reconstruction and novel view synthesis. However, t…
Chemistry-Inspired Diffusion with Non-Differentiable Guidance
Yuchen Shen, Chenhao Zhang, Sijie Fu +3
Recent advances in diffusion models have shown remarkable potential in the conditional generation of novel molecules. These models can be guided in two ways: (i) explicitly, throug…