5 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CV2023
Guiding Image Captioning Models Toward More Specific Captions
Simon Kornblith, Lala Li, Zirui Wang +1
Image captioning is conventionally formulated as the task of generating captions for images that match the distribution of reference image-caption pairs. However, reference caption…
cs.LG2023★ 5 cited
FIT: Far-reaching Interleaved Transformers
Ting Chen, Lala Li
We present FIT: a transformer-based architecture with efficient self-attention and adaptive computation. Unlike original transformers, which operate on a single sequence of data to…