3 papers
cs.CV2026
Cost-Aware Routing for Efficient Text-To-Image Generation
Qinchan Li, Kenneth Chen, Changyue Su +3
Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity als…
cs.CV2026
Infinite Gaze Generation for Videos with Autoregressive Diffusion
Jenna Kang, Colin Groth, Tong Wu +4
Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions…
cs.CV2025
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
Jenna Kang, Maria Silva, Patsorn Sangkloy +3
Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their…