papers

Publications (15)

cs.AI2025

Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem

Declan Campbell, Sunayana Rane, Tyler Giallanza +8

Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image…

cs.CV2019

Transferable Representation Learning in Vision-and-Language Navigation

Haoshuo Huang, Vihan Jain, Harsh Mehta +4

Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic envir…

cs.CV2022

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Jiahui Yu, Yuanzhong Xu, Jing Yu Koh +14

We present the Pathways Autoregressive Text-to-Image (Parti) model, which generates high-fidelity photorealistic images and supports content-rich synthesis involving complex compos…

cs.CV2022

Vector-quantized Image Modeling with Improved VQGAN

Jiahui Yu, Xin Li, Jing Yu Koh +7

Pretraining language models with next-token prediction on massive text corpora has delivered phenomenal zero-shot, few-shot, transfer learning and multi-tasking capabilities on bot…

cs.AI2021

On the Evaluation of Vision-and-Language Navigation Instructions

Ming Zhao, Peter Anderson, Vihan Jain +4

Vision-and-Language Navigation wayfinding agents can be enhanced by exploiting automatically generated navigation instructions. However, existing instruction generators have not be…

cs.MA2026

Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs

Elizabeth Mieczkowski, Alexander Ku, Tiwalayo Eisape +5

Large language models (LLMs) are increasingly deployed in teams, yet existing coordination approaches often occupy two extremes. Highly structured methods rely on fixed roles, pipe…

cs.CV2023

Prompt Expansion for Adaptive Text-to-Image Generation

Siddhartha Datta, Alexander Ku, Deepak Ramachandran +1

Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the images can be repetitive. This paper proposes a Pro…

cs.LG2023

A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning

Aishwarya Kamath, Peter Anderson, Su Wang +6

Recent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots…

cs.RO2019

General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping

Gabriel Ilharco, Vihan Jain, Alexander Ku +2

In instruction conditioned navigation, agents interpret natural language and their surroundings to navigate through an environment. Datasets for studying this task typically contai…

cs.CV2018

Image Transformer

Niki Parmar, Ashish Vaswani, Jakob Uszkoreit +4

Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of mod…

cs.CV2024

DOCCI: Descriptions of Connected and Contrasting Images

Yasumasa Onoe, Sunayana Rane, Zachary Berger +9

Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would al…

cs.CV2020

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Alexander Ku, Peter Anderson, Roma Patel +2

We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) t…

cs.LG2023

Gaussian Process Probes (GPP) for Uncertainty-Aware Probing

Zi Wang, Alexander Ku, Jason Baldridge +2

Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data.…

cs.AI2019

Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation

Vihan Jain, Gabriel Magalhaes, Alexander Ku +3

Advances in learning and representations have reinvigorated work that connects language to other modalities. A particularly exciting direction is Vision-and-Language Navigation(VLN…

cs.CV2021

PanGEA: The Panoramic Graph Environment Annotation Toolkit

Alexander Ku, Peter Anderson, Jordi Pont-Tuset +1

PanGEA, the Panoramic Graph Environment Annotation toolkit, is a lightweight toolkit for collecting speech and text annotations in photo-realistic 3D environments. PanGEA immerses…