2 citations · 4 across the 5 of their papers we have counts for
3 papers · 2 filters
UniMo: Unifying Human and Animal Motion Generation
Zeyu Zhang, Zhiyuan Zhang, Siheng Wang +4
The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending r…
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad +8
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on…
GeoWorld: Geometric World Models
Zeyu Zhang, Danning Li, Ian Reid +1
Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, e…