most citedTS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval

9 citations · 13 across the 9 of their papers we have counts for

collaborators

8 papers

cs.CV2023

InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation

Anwen Hu, Shizhe Chen, Liang Zhang +1

Automatic image captioning evaluation is critical for benchmarking and promoting advances in image captioning research. Existing metrics only provide a single score to measure capt…

cs.CV20231 cited

Knowledge Enhanced Model for Live Video Comment Generation

Jieting Chen, Junkai Ding, Wenping Chen +1

Live video commenting is popular on video media platforms, as it can create a chatting atmosphere and provide supplementary information for users while watching videos. Automatical…

cs.CL2023

MPMQA: Multimodal Question Answering on Product Manuals

Liang Zhang, Anwen Hu, Jing Zhang +2

Visual contents, such as illustrations and images, play a big role in product manual understanding. Existing Product Manual Question Answering (PMQA) datasets tend to ignore visual…

cs.SD2023

PHONEix: Acoustic Feature Processing Strategy for Enhanced Singing Pronunciation with Phoneme Distribution Predictor

Yuning Wu, Jiatong Shi, Tao Qian +2

Singing voice synthesis (SVS), as a specific task for generating the vocal singing voice from a music score, has drawn much attention in recent years. SVS faces the challenge that…

cs.CV20231 cited

Accommodating Audio Modality in CLIP for Multimodal Processing

Ludan Ruan, Anwen Hu, Yuqing Song +3

Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training,…

cs.CV20222 cited

Exploring Anchor-based Detection for Ego4D Natural Language Query

Sipeng Zheng, Qi Zhang, Bei Liu +2

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehen…