1 paper
Pin Ji, Yang Feng, Weitao Huang +2
The development of modern NLP applications often relies on various benchmark datasets containing plenty of manually labeled tests to evaluate performance. While constructing datase…