3 papers
cs.CV2026
GeoWorld: Geometric World Models
Zeyu Zhang, Danning Li, Ian Reid +1
Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, e…
cs.CV2026
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad +8
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on…
cs.CV2025
Motion Anything: Any to Motion Generation
Zeyu Zhang, Yiran Wang, Wei Mao +7
Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperfo…