185 citations · 539 across the 27 of their papers we have counts for
3 papers · 1 filter
Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment
Lu Yu, Malvina Nikandrou, Jiali Jin +1
Automated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and e…
Going for GOAL: A Resource for Grounded Football Commentaries
Alessandro Suglia, José Lopes, Emanuele Bastianelli +6
Recent video+language datasets cover domains where the interaction is highly structured, such as instructional videos, or where the interaction is scripted, such as TV shows. Both…
History for Visual Dialog: Do we really need it?
Shubham Agarwal, Trung Bui, Joon-Young Lee +2
Visual Dialog involves "understanding" the dialog history (what has been discussed previously) and the current question (what is asked), in addition to grounding information in the…