26 citations · 27 across the 2 of their papers we have counts for
4 papers
ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
Zecheng He, Srinivas Sunkara, Xiaoxue Zang +7
As mobile devices are becoming ubiquitous, regularly interacting with a variety of user interfaces (UIs) is a common aspect of daily life for many people. To improve the accessibil…
MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines
Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara +3
MultiWOZ is a well-known task-oriented dialogue dataset containing over 10,000 annotated dialogues spanning 8 domains. It is extensively used as a benchmark for dialogue state trac…
Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz +3
Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question…
Resolving Referring Expressions in Images With Labeled Elements
Nevan Wichers, Dilek Hakkani-Tur, Jindong Chen
Images may have elements containing text and a bounding box associated with them, for example, text identified via optical character recognition on a computer screen image, or a na…