88 citations · 153 across the 22 of their papers we have counts for
7 papers · 1 filter
GPA: Learning GUI Process Automation from Demonstrations
Zirui Zhao, Jun Hao Liew, Yan Yang +5
GUI Process Automation (GPA) is a lightweight but general vision-based Robotic Process Automation (RPA), which enables fast and stable process replay with only a single demo. Addre…
Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions
David Junhao Zhang, Dongxu Li, Hung Le +3
Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the g…
BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues
Hung Le, Doyen Sahoo, Nancy F. Chen +1
Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances whi…
Adaptive Task Sampling for Meta-Learning
Chenghao Liu, Zhihao Wang, Doyen Sahoo +3
Meta-learning methods have been extensively studied and applied in computer vision, especially for few-shot classification tasks. The key idea of meta-learning for few-shot classif…
FoodAI: Food Image Recognition via Deep Learning for Smart Food Logging
Doyen Sahoo, Wang Hao, Shu Ke +5
An important aspect of health monitoring is effective logging of food consumption. This can help management of diet-related diseases like obesity, diabetes, and even cardiovascular…
Recent Advances in Deep Learning for Object Detection
Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi
Object detection is a fundamental visual recognition problem in computer vision and has been widely studied in the past decades. Visual object detection aims to find objects of cer…