3 papers
cs.CV2021
Video Question Answering with Phrases via Semantic Roles
Arka Sadhu, Kan Chen, Ram Nevatia
Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA model…
cs.CV2020
Utilizing Every Image Object for Semi-supervised Phrase Grounding
Haidong Zhu, Arka Sadhu, Zhaoheng Zheng +1
Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the va…
cs.CV2020
Curriculum DeepSDF
Yueqi Duan, Haidong Zhu, He Wang +3
When learning to sketch, beginners start with simple and flexible shapes, and then gradually strive for more complex and accurate ones in the subsequent training sessions. In this…