2 papers
cs.CV2021
MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants
Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen +3
In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. C…
cs.CV2020
Generating Natural Questions from Images for Multimodal Assistants
Alkesh Patel, Akanksha Bindal, Hadas Kotek +2
Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in…