activity
20172026
most citedXGPT: Cross-modal Generative Pre-Training for Image Captioning

20 citations · 20 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang +7

Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either s…

cs.CL2020

Multi-View Learning for Vision-and-Language Navigation

Qiaolin Xia, Xiujun Li, Chunyuan Li +5

Learning to navigate in a visual environment following natural language instructions is a challenging task because natural language instructions are highly variable, ambiguous, and…

cs.CL202020 cited

XGPT: Cross-modal Generative Pre-Training for Image Captioning

Qiaolin Xia, Haoyang Huang, Nan Duan +7

While many BERT-based cross-modal pre-trained models produce excellent results on downstream understanding tasks like image-text retrieval and VQA, they cannot be applied to genera…

cs.CL2019

Robust Navigation with Language Pretraining and Stochastic Sampling

Xiujun Li, Chunyuan Li, Qiaolin Xia +5

Core to the vision-and-language navigation (VLN) challenge is building robust instruction representations and action decoding schemes, which can generalize well to previously unsee…

cs.CL2018

Incorporating Glosses into Neural Word Sense Disambiguation

Fuli Luo, Tianyu Liu, Qiaolin Xia +2

Word Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context. Lexical resources like WordNet which are proved to be of great h…

cs.CL2017

Improving Chinese SRL with Heterogeneous Annotations

Qiaolin Xia, Baobao Chang, Zhifang Sui

Previous studies on Chinese semantic role labeling (SRL) have concentrated on single semantically annotated corpus. But the training data of single corpus is often limited. Meanwhi…