10 papers
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
Bingxuan Li, Jiahao Wu, Yuan Xu +6
Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introdu…
Configurable Holography: Towards Display and Scene Adaptation
Yicheng Zhan, Liang Shi, Wojciech Matusik +2
Rendering holograms for holographic displays is often an iterative and computationally costly process. Emerging learned holography methods have alleviated this bottleneck by enabli…
Cost-Aware Routing for Efficient Text-To-Image Generation
Qinchan Li, Kenneth Chen, Changyue Su +3
Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity als…
Infinite Gaze Generation for Videos with Autoregressive Diffusion
Jenna Kang, Colin Groth, Tong Wu +4
Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions…
HOICraft: In-Situ VLM-based Authoring Tool for Part-Level Hand-Object Interaction Design in VR
Dohui Lee, Qi Sun, Sang Ho Yoon
Hand-Object Interaction (HOI) is a key interaction component in Virtual Reality (VR). However, designing HOI still requires manual efforts to decide how object should be selected a…
Foveation Improves Payload Capacity in Steganography
Lifeng Qiu Lin, Henry Kam, Qi Sun +1
Steganography finds its use in visual medium such as providing metadata and watermarking. With support of efficient latent representations and foveated rendering, we trained models…