157 citations · 172 across the 3 of their papers we have counts for
7 papers
Infinite Gaze Generation for Videos with Autoregressive Diffusion
Jenna Kang, Colin Groth, Tong Wu +4
Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions…
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
Jenna Kang, Maria Silva, Patsorn Sangkloy +3
Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their…
Cost-Aware Routing for Efficient Text-To-Image Generation
Qinchan Li, Kenneth Chen, Changyue Su +3
Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity als…
Argoverse: 3D Tracking and Forecasting with Rich Maps
Ming-Fang Chang, John Lambert, Patsorn Sangkloy +8
We present Argoverse -- two datasets designed to support autonomous vehicle machine learning tasks such as 3D tracking and motion forecasting. Argoverse was collected by a fleet of…
Kernel Mean Matching for Content Addressability of GANs
Wittawat Jitkrittum, Patsorn Sangkloy, Muhammad Waleed Gondal +3
We propose a novel procedure which adds "content-addressability" to any given unconditional implicit model e.g., a generative adversarial network (GAN). The procedure allows users…
Informative Features for Model Comparison
Wittawat Jitkrittum, Heishiro Kanagawa, Patsorn Sangkloy +3
Given two candidate models, and a set of target observations, we address the problem of measuring the relative goodness of fit of the two models. We propose two new statistical tes…