24 citations · 29 across the 6 of their papers we have counts for
6 papers
DiT: Efficient Vision Transformers with Dynamic Token Routing
Yuchen Ma, Zhengcong Fei, Junshi Huang
Recently, the tokens of images share the same static data flow in many dense networks. However, challenges arise from the variance among the objects in images, such as large variat…
Divide and Adapt: Active Domain Adaptation via Customized Learning
Duojun Huang, Jichang Li, Weikai Chen +3
Active domain adaptation (ADA) aims to improve the model adaptation performance by incorporating active learning (AL) techniques to label a maximally-informative subset of target s…
Gradient-Free Textual Inversion
Zhengcong Fei, Mingyuan Fan, Junshi Huang
Recent works on personalized text-to-image generation usually learn to bind a special token with specific subjects or styles of a few given images by tuning its embedding through g…
EfficientRep:An Efficient Repvgg-style ConvNets with Hardware-aware Neural Network Design
Kaiheng Weng, Xiangxiang Chu, Xiaoming Xu +2
We present a hardware-efficient architecture of convolutional neural network, which has a repvgg-like architecture. Flops or parameters are traditional metrics to evaluate the effi…
PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding
Zihan Ding, Zi-han Ding, Tianrui Hui +4
Panoptic Narrative Grounding (PNG) is an emerging task whose goal is to segment visual objects of things and stuff categories described by dense narrative captions of a still image…
Efficient Modeling of Future Context for Image Captioning
Zhengcong Fei, Junshi Huang, Xiaoming Wei +1
Existing approaches to image captioning usually generate the sentence word-by-word from left to right, with the constraint of conditioned on local context including the given image…