1 citations · 1 across the 4 of their papers we have counts for
5 papers
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Shenyuan Gao, William Liang, Kaiyuan Zheng +27
Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, espe…
Scalable Policy Evaluation with Video World Models
Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang +4
Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluati…
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
Wei-Cheng Tseng, David Harwath
Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their poten…
Probing the Robustness Properties of Neural Speech Codecs
Wei-Cheng Tseng, David Harwath
Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategi…
Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present th…