55 citations · 107 across the 12 of their papers we have counts for
14 papers
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
Georgia Gabriela Sampaio, Ruixiang Zhang, Shuangfei Zhai +4
Evaluating text-to-image generative models remains a challenge, despite the remarkable progress being made in their overall performances. While existing metrics like CLIPScore work…
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
Chen Huang, Skyler Seto, Samira Abnar +3
Large pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e.g., satellite imagery) or fine-grained…
On the benefits of pixel-based hierarchical policies for task generalization
Tudor Cristea-Platon, Bogdan Mazoure, Josh Susskind +1
Reinforcement learning practitioners often avoid hierarchical policies, especially in image-based observation spaces. Typically, the single-task performance improvement over flat-p…
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
Yuhang Zang, Hanlin Goh, Josh Susskind +1
Existing vision-language models exhibit strong generalization on a variety of visual domains and tasks. However, such models mainly perform zero-shot recognition in a closed-set ma…
What Algorithms can Transformers Learn? A Study in Length Generalization
Hattie Zhou, Arwen Bradley, Etai Littwin +5
Large language models exhibit surprising emergent generalization properties, yet also struggle on many simple reasoning tasks such as arithmetic and parity. This raises the questio…
Adaptivity and Modularity for Efficient Generalization Over Task Complexity
Samira Abnar, Omid Saremi, Laurent Dinh +8
Can transformers generalize efficiently on problems that require dealing with examples with different levels of difficulty? We introduce a new task tailored to assess generalizatio…