3 citations · 3 across the 1 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023
TG-VQA: Ternary Game of Video Question Answering
Hao Li, Peng Jin, Zesen Cheng +5
Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions…
cs.CV2023★ 13 cited
PixMIM: Rethinking Pixel Reconstruction in Masked Image Modeling
Yuan Liu, Songyang Zhang, Jiacheng Chen +2
Masked Image Modeling (MIM) has achieved promising progress with the advent of Masked Autoencoders (MAE) and BEiT. However, subsequent works have complicated the framework with new…
cs.CV2022★ 3 cited
GCFSR: a Generative and Controllable Face Super Resolution Method Without Facial and GAN Priors
Jingwen He, Wu Shi, Kai Chen +2
Face image super resolution (face hallucination) usually relies on facial priors to restore realistic details and preserve identity information. Recent advances can achieve impress…