3 papers
cs.CV2026
Scaling Spatial Intelligence with Multimodal Foundation Models
Zhongang Cai, Ruisi Wang, Chenyang Gu +26
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation m…
cs.CV2025
Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory
Hui Pang, Sunil Hadap, Violetta Shevchenko +2
Augmented reality is revolutionizing beauty industry with virtual try-on (VTO) applications, which empowers users to try a wide variety of products using their phones without the h…
cs.CV2025
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation
Wanqi Yin, Zhongang Cai, Ruisi Wang +12
Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art me…