2 papers
cs.SD2021
Streaming on-device detection of device directed speech from voice and touch-based invocation
Ognjen Rudovic, Akanksha Bindal, Vineet Garg +3
When interacting with smart devices such as mobile phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the devic…
cs.CV2020
Generating Natural Questions from Images for Multimodal Assistants
Alkesh Patel, Akanksha Bindal, Hadas Kotek +2
Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in…