10 papers
GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
Yuxuan Xue, Ruofan Liang, Egor Zakharov +6
Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current method…
CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation
Juyeong Hwang, Seong-Eun Hong, Jinhyun Kim +4
Crowds do not merely move; they decide. Human navigation is inherently contextual: people interpret the meaning of space, social norms, and potential consequences before acting. Si…
DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation
Chengyang Zhao, Uksang Yoo, Arkadeep Narayan Chaudhury +4
Hair care is an essential daily activity, yet it remains inaccessible to individuals with limited mobility and challenging for autonomous robot systems due to the fine-grained phys…
CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
Zhiyi Kuang, Chengan He, Egor Zakharov +6
We present CamLit, the first unified video diffusion model that jointly performs novel view synthesis (NVS) and relighting from a single input image. Given one reference image, a u…
Gaussian Pixel Codec Avatars: A Hybrid Representation for Efficient Rendering
Divam Gupta, Anuj Pahuja, Nemanja Bartolovic +3
We present Gaussian Pixel Codec Avatars (GPiCA), photorealistic head avatars that can be generated from multi-view images and efficiently rendered on mobile devices. GPiCA utilizes…
A Real-world Display Inverse Rendering Dataset
Seokjun Choi, Hoon-Gyu Chung, Yujin Jeon +2
Inverse rendering aims to reconstruct geometry and reflectance from captured images. Display-camera imaging systems offer unique advantages for this task: each pixel can easily fun…