24 citations · 30 across the 3 of their papers we have counts for
3 papers
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense
Zhecan Wang, Haoxuan You, Yicheng He +3
Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve compr…
Uncovering Visually Impaired Gamers' Preferences for Spatial Awareness Tools Within Video Games
Vishnu Nair, Shao-en Ma, Ricardo E. Gonzalez Penuela +6
Sighted players gain spatial awareness within video games through sight and spatial awareness tools (SATs) such as minimaps. Visually impaired players (VIPs), however, must often r…
Query Adaptive Few-Shot Object Detection with Heterogeneous Graph Convolutional Networks
Guangxing Han, Yicheng He, Shiyuan Huang +2
Few-shot object detection (FSOD) aims to detect never-seen objects using few examples. This field sees recent improvement owing to the meta-learning techniques by learning how to m…