1 citations · 1 across the 1 of their papers we have counts for
1 paper
Mina Huh, Fangyuan Xu, Yi-Hao Peng +5
Vision language models can now generate long-form answers to questions about images - long-form visual question answers (LFVQA). We contribute VizWiz-LF, a dataset of long-form ans…