most citedANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos

1 citations · 2 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV20241 cited

TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model

Jiahao Lyu, Jin Wei, Gangyan Zeng +4

Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene…

cs.MM2023

Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks

Hongye Liu, Xianhai Xie, Yang Gao +2

The pretrain-then-finetune paradigm has been widely used in various unimodal and multimodal tasks. However, finetuning all the parameters of a pre-trained model becomes prohibitive…

cs.CL2023

Non-Sequential Graph Script Induction via Multimedia Grounding

Yu Zhou, Sha Li, Manling Li +4

Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts…

eess.AS2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

Xiao-Min Zeng, Yan Song, Zhu Zhuo +5

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE)…

cs.CL2023

E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation

Cong Ma, Yaping Zhang, Mei Tu +3

Text image machine translation (TIMT) aims to translate texts embedded in images from one source language to another target language. Existing methods, both two-stage cascade and o…

cs.CV20231 cited

ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos

Zhou Yu, Lixiang Zheng, Zhou Zhao +4

Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…