3 papers
cs.CV2025
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Size Wu, Zhonghua Wu, Zerui Gong +5
In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in…
cs.CV2025
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
Size Wu, Wenwei Zhang, Lumin Xu +6
Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations…
cs.CV2023
Towards Robust and Expressive Whole-body Human Pose and Shape Estimation
Hui EnPang, Zhongang Cai, Lei Yang +4
Whole-body pose and shape estimation aims to jointly predict different behaviors (e.g., pose, hand gesture, facial expression) of the entire human body from a monocular image. Exis…