activity
20222024
most citedVLMAE: Vision-Language Masked Autoencoder

9 citations · 17 across the 5 of their papers we have counts for

collaborators

5 papers

cs.IR20245 cited

RAT: Retrieval-Augmented Transformer for Click-Through Rate Prediction

Yushen Li, Jinpeng Wang, Tao Dai +4

Predicting click-through rates (CTR) is a fundamental task for Web applications, where a key issue is to devise effective models for feature interactions. Current methodologies pre…

cs.CV20231 cited

Towards Robust Scene Text Image Super-resolution via Explicit Location Enhancement

Hang Guo, Tao Dai, Guanghao Meng +1

Scene text image super-resolution (STISR), aiming to improve image quality while boosting downstream scene text recognition accuracy, has recently achieved great success. However,…

cs.CV20221 cited

Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature Domain

Yujun Huang, Bin Chen, Shiyu Qin +4

Beyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., an…

cs.CV20229 cited

VLMAE: Vision-Language Masked Autoencoder

Sunan He, Taian Guo, Tao Dai +4

Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…

eess.IV20221 cited

Adaptive Local Implicit Image Function for Arbitrary-scale Super-resolution

Hongwei Li, Tao Dai, Yiming Li +2

Image representation is critical for many visual tasks. Instead of representing images discretely with 2D arrays of pixels, a recent study, namely local implicit image function (LI…