2 papers
cs.CL2020
NLPStatTest: A Toolkit for Comparing NLP System Performance
Haotian Zhu, Denise Mak, Jesse Gioannini +1
Statistical significance testing centered on p-values is commonly used to compare NLP system performance, but p-values alone are insufficient because statistical significance diffe…
cs.CL2020
Probing for Multilingual Numerical Understanding in Transformer-Based Language Models
Devin Johnson, Denise Mak, Drew Barker +1
Natural language numbers are an example of compositional structures, where larger numbers are composed of operations on smaller numbers. Given that compositional reasoning is a key…