33 citations · 44 across the 9 of their papers we have counts for
9 papers
OmniZoomer: Learning to Move and Zoom in on Sphere at High-Resolution
Zidong Cao, Hao Ai, Yan-Pei Cao +3
Omnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive en…
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video
Jia-Wei Liu, Yan-Pei Cao, Tianyuan Yang +5
We introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video…
MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
Mingdeng Cao, Xintao Wang, Zhongang Qi +3
Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results…
VMesh: Hybrid Volume-Mesh Representation for Efficient View Synthesis
Yuan-Chen Guo, Yan-Pei Cao, Chen Wang +4
With the emergence of neural radiance fields (NeRFs), view synthesis quality has reached an unprecedented level. Compared to traditional mesh-based assets, this volumetric represen…
Accelerating Vision-Language Pretraining with Free Language Modeling
Teng Wang, Yixiao Ge, Feng Zheng +4
The state of the arts in vision-language pretraining (VLP) achieves exemplary performance but suffers from high training costs resulting from slow convergence and long training tim…
T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
Chong Mou, Xintao Wang, Liangbin Xie +5
The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying…