14 papers
GeoStereo: A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal
Qizhe Wei, Xianda Guo, Shaocong Xu +3
Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in ch…
Relit-LiVE: Relight Video by Jointly Learning Environment Video
Weiqing Xiao, Hong Li, Xiuyu Yang +7
Recent advances have shown that large-scale video diffusion models can be repurposed as neural renderers by first decomposing videos into intrinsic scene representations and then p…
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Houyuan Chen, Hong Li, Xianghao Kong +8
Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each…
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
Mingju Gao, Kaisen Yang, Huan-ang Gao +13
Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains frag…
NeAR: Coupled Neural Asset-Renderer Stack
Hong Li, Chongjie Ye, Houyuan Chen +12
Neural asset authoring and neural rendering have traditionally evolved as disjoint paradigms: one generates digital assets for fixed graphics pipelines, while the other maps conven…
Light of Normals: Unified Feature Representation for Universal Photometric Stereo
Houyuan Chen, Hong Li, Chongjie Ye +11
Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination model…