2 citations · 2 across the 2 of their papers we have counts for
4 papers · 1 filter
Exploring Out-of-Distribution Generalization in Text Classifiers Trained on Tobacco-3482 and RVL-CDIP
Stefan Larson, Navtej Singh, Saarthak Maheshwari +2
To be robust enough for widespread adoption, document analysis systems involving machine learning models must be able to respond correctly to inputs that fall outside of the data d…
LSOIE: A Large-Scale Dataset for Supervised Open Information Extraction
Jacob Solawetz, Stefan Larson
Open Information Extraction (OIE) systems seek to compress the factual propositions of a sentence into a series of n-ary tuples. These tuples are useful for downstream tasks in nat…
An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction
Stefan Larson, Anish Mahendran, Joseph J. Peper +8
Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover eve…
Outlier Detection for Improved Data Quality and Diversity in Dialog Systems
Stefan Larson, Anish Mahendran, Andrew Lee +6
In a corpus of data, outliers are either errors: mistakes in the data that are counterproductive, or are unique: informative samples that improve model robustness. Identifying outl…