2 citations · 2 across the 9 of their papers we have counts for
7 papers · 1 filter
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
David Yifan Yao, Albert J. Zhai, Shenlong Wang
This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth predictio…
AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
Hao-Yu Hsu, Zhi-Hao Lin, Albert Zhai +2
Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and…
CropCraft: Complete Structural Characterization of Crop Plants From Images
Albert J. Zhai, Xinlei Wang, Kaiyuan Li +6
The ability to automatically build 3D digital twins of plants from images has countless applications in agriculture, environmental science, robotics, and other fields. However, cur…
EarthGen: Generating the World from Top-Down Views
Ansh Sharma, Albert Xiao, Praneet Rathi +4
In this work, we present a novel method for extensive multi-scale generative terrain modeling. At the core of our model is a cascade of superresolution diffusion models that can be…
Physical Property Understanding from Language-Embedded Feature Fields
Albert J. Zhai, Yuan Shen, Emily Y. Chen +5
Can computers perceive the physical properties of objects solely through vision? Research in cognitive science and vision science has shown that humans excel at identifying materia…
On the Overconfidence Problem in Semantic 3D Mapping
Joao Marcos Correia Marques, Albert Zhai, Shenlong Wang +1
Semantic 3D mapping, the process of fusing depth and image segmentation information between multiple views to build 3D maps annotated with object classes in real-time, is a recent…