12 citations · 17 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 12 cited
ViperGPT: Visual Inference via Python Execution for Reasoning
Dídac Surís, Sachit Menon, Carl Vondrick
Answering visual queries is a complex task that requires both visual processing and reasoning. End-to-end models, the dominant approach for this task, do not explicitly differentia…
cs.CV2023★ 3 cited
Affective Faces for Goal-Driven Dyadic Communication
Scott Geng, Revant Teotia, Purva Tendulkar +2
We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approac…
cs.LG2022★ 2 cited
Forget-me-not! Contrastive Critics for Mitigating Posterior Collapse
Sachit Menon, David Blei, Carl Vondrick
Variational autoencoders (VAEs) suffer from posterior collapse, where the powerful neural networks used for modeling and inference optimize the objective without meaningfully using…