14 citations · 21 across the 5 of their papers we have counts for
7 papers
TPA-Net: Generate A Dataset for Text to Physics-based Animation
Yuxing Qiu, Feng Gao, Minchen Li +3
Recent breakthroughs in Vision-Language (V&L) joint research have achieved remarkable results in various text-driven tasks. High-quality Text-to-video (T2V), a task that has been l…
Towards Reasoning-Aware Explainable VQA
Rakesh Vaideeswaran, Feng Gao, Abhinav Mathur +1
The domain of joint vision-language understanding, especially in the context of reasoning in Visual Question Answering (VQA) models, has garnered significant attention in the recen…
A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai +3
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information…
A Comparative Review of Recent Few-Shot Object Detection Algorithms
Leng Jiaxu, Chen Taiyue, Gao Xinbo +4
Few-shot object detection, learning to adapt to the novel classes with a few labeled data, is an imperative and long-lasting problem due to the inherent long-tail distribution of r…
Towards Efficient Full 8-bit Integer DNN Online Training on Resource-limited Devices without Batch Normalization
Yukuan Yang, Xiaowei Chi, Lei Deng +3
Huge computational costs brought by convolution and batch normalization (BN) have caused great challenges for the online training and corresponding applications of deep neural netw…
Capability Iteration Network for Robot Path Planning
Buqing Nie, Yue Gao, Yidong Mei +1
Path planning is an important topic in robotics. Recently, value iteration based deep learning models have achieved good performance such as Value Iteration Network(VIN). However,…