1 paper · 1 filter
Clayton Fields, Casey Kennington
Vision language tasks, such as answering questions about or generating captions that describe an image, are difficult tasks for computers to perform. A relatively recent body of re…