1 citations · 2 across the 8 of their papers we have counts for
8 papers
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
Jiahao Lyu, Jin Wei, Gangyan Zeng +4
Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene…
Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks
Hongye Liu, Xianhai Xie, Yang Gao +2
The pretrain-then-finetune paradigm has been widely used in various unimodal and multimodal tasks. However, finetuning all the parameters of a pre-trained model becomes prohibitive…
Non-Sequential Graph Script Induction via Multimedia Grounding
Yu Zhou, Sha Li, Manling Li +4
Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts…
Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
Xiao-Min Zeng, Yan Song, Zhu Zhuo +5
In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE)…
E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma, Yaping Zhang, Mei Tu +3
Text image machine translation (TIMT) aims to translate texts embedded in images from one source language to another target language. Existing methods, both two-stage cascade and o…
ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos
Zhou Yu, Lixiang Zheng, Zhou Zhao +4
Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…