1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.RO2025
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8
While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…
eess.AS2024★ 1 cited
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
Kevin Cai, Chonghua Liu, David M. Chan
The Internet's wealth of content, with up to 60% published in English, starkly contrasts the global population, where only 18.8% are English speakers, and just 5.1% consider it the…