5 papers
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
Bingxuan Li, Jiahao Wu, Yuan Xu +6
Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introdu…
Cost-Aware Routing for Efficient Text-To-Image Generation
Qinchan Li, Kenneth Chen, Changyue Su +3
Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity als…
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
Jenna Kang, Maria Silva, Patsorn Sangkloy +3
Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their…
Image-GS: Content-Adaptive Image Representation via 2D Gaussians
Yunxiang Zhang, Bingxuan Li, Alexandr Kuznetsov +6
Neural image representations have emerged as a promising approach for encoding and rendering visual data. Combined with learning-based workflows, they demonstrate impressive trade-…
BudgetFusion: Perceptually-Guided Adaptive Diffusion Models
Qinchan Li, Kenneth Chen, Changyue Su +1
Diffusion models have shown unprecedented success in the task of text-to-image generation. While these models are capable of generating high-quality and realistic images, the compl…