10 citations · 17 across the 10 of their papers we have counts for
14 papers · 1 filter
ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation
Jiangtao Kong, Peijun Zhao, Chun-Fu Chen +4
Incremental Learning (IL) for Open-ended Image-to-Text Generation (OpenITG) enables models to continuously generate accurate, contextually relevant text for new images while preser…
Temporal Relevance Analysis for Video Action Models
Quanfu Fan, Donghyun Kim, Chun-Fu +4
In this paper, we provide a deep analysis of temporal modeling for action recognition, an important but underexplored problem in the literature. We first propose a new approach to…
Dynamic Network Quantization for Efficient Video Inference
Ximeng Sun, Rameswar Panda, Chun-Fu Chen +3
Deep convolutional networks have recently achieved great success in video recognition, yet their practical realization remains a challenge due to the large amount of computational…
Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled Data
Ashraful Islam, Chun-Fu Chen, Rameswar Panda +3
Most existing works in few-shot learning rely on meta-learning the network on a large base dataset which is typically from the same domain as the target dataset. We tackle the prob…
AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition
Rameswar Panda, Chun-Fu Chen, Quanfu Fan +4
Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learni…
Detector-Free Weakly Supervised Grounding by Separation
Assaf Arbelle, Sivan Doveh, Amit Alfassy +14
Nowadays, there is an abundance of data involving images and surrounding free-form text weakly corresponding to those images. Weakly Supervised phrase-Grounding (WSG) deals with th…