9 papers
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
Weimin Bai, Yubo Li, Weijian Luo +4
Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations…
LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)
Wei Luo, Yiting Lu, Xin Li +32
This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The…
Multimodal OCR: Parse Anything from Documents
Handong Zheng, Yumeng Li, Kaile Zhang +22
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…
Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
Yuxuan Gu, Weimin Bai, Yifei Wang +2
Masked auto-regressive diffusion models (MAR) benefit from the expressive modeling ability of diffusion models and the flexibility of masked auto-regressive ordering. However, vani…
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
Zihan Ye, Shreyank N Gowda, Kaile Du +2
Zero-shot Learning (ZSL) enables classifiers to recognize classes unseen during training, commonly via generative two stage methods: (1) learn visual semantic correlations from see…
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
Yifei Wang, Weimin Bai, Colin Zhang +3
In this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, -distill, etc, inside a theory-driven framework w…