7 citations · 17 across the 13 of their papers we have counts for
6 papers · 1 filter
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
Yue Han, Chong Li, Zhening Liu +5
Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due to the scarcity of large-scal…
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
Dong Chen, Fangyun Wei, Ziyu Wan +18
We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters acro…
Spatia: Video Generation with Updatable Spatial Memory
Jinjing Zhao, Fangyun Wei, Zhening Liu +3
Existing video generation models struggle to maintain long-term spatial and temporal consistency due to the dense, high-dimensional nature of video signals. To overcome this limita…
CustomX: Unified Character, Action, and Scene Customization in Video World Models
Yitong Wang, Fangyun Wei, Hongyang Zhang +2
Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, whic…
From Virtual Games to Real-World Play
Wenqiang Sun, Fangyun Wei, Jinjing Zhao +5
We introduce RealPlay, a neural network-based real-world game engine that enables interactive video generation from user control signals. Unlike prior works focused on game-style v…
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
Jierun Chen, Fangyun Wei, Jinjing Zhao +5
Referring expression comprehension (REC) involves localizing a target instance based on a textual description. Recent advancements in REC have been driven by large multimodal model…