21 citations · 40 across the 4 of their papers we have counts for
7 papers
RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation
Md Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh +3
Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained…
Knowledge Distillation with Noisy Labels for Natural Language Understanding
Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid +5
Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of…
Context-aware Adversarial Training for Name Regularity Bias in Named Entity Recognition
Abbas Ghaddar, Philippe Langlais, Ahmad Rashid +1
In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed…
WiRe57 : A Fine-Grained Benchmark for Open Information Extraction
William Léchelle, Fabrizio Gotti, Philippe Langlais
We build a reference for the task of Open Information Extraction, on five documents. We tentatively resolve a number of issues that arise, including inference and granularity. We s…
Robust Lexical Features for Improved Neural Network Named-Entity Recognition
Abbas Ghaddar, Philippe Langlais
Neural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical featur…
Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
Francis Grégoire, Philippe Langlais
Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose a bidirectional recurrent neu…