1 paper
Peter Hase, Mohit Bansal, Peter Clark +1
How can we train models to perform well on hard test data when hard training data is by definition difficult to label correctly? This question has been termed the scalable oversigh…