activity
20242026
collaborators

5 papers

physics.optics2026

Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding

Bingxuan Li, Jiahao Wu, Yuan Xu +6

Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introdu…

cs.CV2026

Cost-Aware Routing for Efficient Text-To-Image Generation

Qinchan Li, Kenneth Chen, Changyue Su +3

Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity als…

cs.CV2025

GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts

Jenna Kang, Maria Silva, Patsorn Sangkloy +3

Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their…

cs.CV2025

Image-GS: Content-Adaptive Image Representation via 2D Gaussians

Yunxiang Zhang, Bingxuan Li, Alexandr Kuznetsov +6

Neural image representations have emerged as a promising approach for encoding and rendering visual data. Combined with learning-based workflows, they demonstrate impressive trade-…

cs.CV2024

BudgetFusion: Perceptually-Guided Adaptive Diffusion Models

Qinchan Li, Kenneth Chen, Changyue Su +1

Diffusion models have shown unprecedented success in the task of text-to-image generation. While these models are capable of generating high-quality and realistic images, the compl…