630 citations · 2.3k across the 46 of their papers we have counts for
5 papers · 1 filter
Voxel-informed Language Grounding
Rodolfo Corona, Shizhan Zhu, Dan Klein +1
Natural language applied to natural 2D images describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3…
Discovering Non-monotonic Autoregressive Orderings with Variational Inference
Xuanlin Li, Brandon Trabucco, Dong Huk Park +4
The predominant approach for language modeling is to process sequences from left to right, but this eliminates a source of information: the order by which the sequence was generate…
Modular Networks for Compositional Instruction Following
Rodolfo Corona, Daniel Fried, Coline Devin +2
Standard architectures used in instruction following often struggle on novel compositions of subgoals (e.g. navigating to landmarks or picking up objects) observed during training.…
Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
Ronghang Hu, Daniel Fried, Anna Rohrbach +3
Vision-and-Language Navigation (VLN) requires grounding instructions, such as "turn right and stop at the door", to routes in a visual environment. The actual grounding can connect…
Object Hallucination in Image Captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns +2
Despite continuously improving performance, contemporary image captioning models are prone to "hallucinating" objects that are not actually in a scene. One problem is that standard…