1.6k citations · 1.9k across the 15 of their papers we have counts for
1 paper · 2 filters
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin +1
Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches fo…