2 papers
cs.CL2022
Understanding and Improving Knowledge Distillation for Quantization-Aware Training of Large Transformer Encoders
Minsoo Kim, Sihwa Lee, Sukjin Hong +2
Knowledge distillation (KD) has been a ubiquitous method for model compression to strengthen the capability of a lightweight model with the transferred knowledge from the teacher.…
cs.CL2020
Tackling Domain-Specific Winograd Schemas with Knowledge-Based Reasoning and Machine Learning
Suk Joon Hong, Brandon Bennett
The Winograd Schema Challenge (WSC) is a common-sense reasoning task that requires background knowledge. In this paper, we contribute to tackling WSC in four ways. Firstly, we sugg…