activity
20182022
most citedTowards Robust Visual Question Answering: Making the Most of Biased Samples via Contrastive Learning

5 citations · 14 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CL20222 cited

COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language Models

Bowen Shen, Zheng Lin, Yuanxin Liu +3

Transformer-based pre-trained language models (PLMs) mostly suffer from excessive overhead despite their advanced capacity. For resource-constrained devices, there is an urgent nee…

cs.CL20222 cited

A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models

Yuanxin Liu, Fandong Meng, Zheng Lin +5

Despite the remarkable success of pre-trained language models (PLMs), they still face two challenges: First, large-scale PLMs are inefficient in terms of memory footprint and compu…

cs.CV20222 cited

Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA

Qingyi Si, Fandong Meng, Mingyu Zheng +6

Visual Question Answering (VQA) models are prone to learn the shortcut solution formed by dataset biases rather than the intended solution. To evaluate the VQA models' reasoning ab…

cs.CV20225 cited

Towards Robust Visual Question Answering: Making the Most of Biased Samples via Contrastive Learning

Qingyi Si, Yuanxin Liu, Fandong Meng +5

Models for Visual Question Answering (VQA) often rely on the spurious correlations, i.e., the language priors, that appear in the biased samples of training set, which make them br…

cs.CL2022

Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training

Yuanxin Liu, Fandong Meng, Zheng Lin +4

Recent studies on the lottery ticket hypothesis (LTH) show that pre-trained language models (PLMs) like BERT contain matching subnetworks that have similar transfer learning perfor…

cs.CL2021

Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation

Yuanxin Liu, Fandong Meng, Zheng Lin +2

Recently, knowledge distillation (KD) has shown great success in BERT compression. Instead of only learning from the teacher's soft label as in conventional KD, researchers find th…