6 papers
Learning from Convolution-based Unlearnable Datasets
Dohyun Kim, Pedro Sandoval-Segura
The construction of large datasets for deep learning has raised concerns regarding unauthorized use of online data, leading to increased interest in protecting data from third-part…
AutoProtoNet: Interpretability for Prototypical Networks
Pedro Sandoval-Segura, Wallace Lawson
In meta-learning approaches, it is difficult for a practitioner to make sense of what kind of representations the model employs. Without this ability, it can be difficult to both u…
An Information-Theoretic Perspective on Overfitting and Underfitting
Daniel Bashir, George D. Montanez, Sonia Sehra +2
We present an information-theoretic framework for understanding overfitting and underfitting in machine learning and prove the formal undecidability of determining whether an arbit…
The Labeling Distribution Matrix (LDM): A Tool for Estimating Machine Learning Algorithm Capacity
Pedro Sandoval Segura, Julius Lauw, Daniel Bashir +4
Algorithm performance in supervised learning is a combination of memorization, generalization, and luck. By estimating how much information an algorithm can memorize from a dataset…
Harvey Mudd College at SemEval-2019 Task 4: The Clint Buchanan Hyperpartisan News Detector
Mehdi Drissi, Pedro Sandoval, Vivaswat Ojha +1
We investigate the recently developed Bidirectional Encoder Representations from Transformers (BERT) model for the hyperpartisan news detection task. Using a subset of hand-labeled…
Program Language Translation Using a Grammar-Driven Tree-to-Tree Model
Mehdi Drissi, Olivia Watkins, Aditya Khant +5
The task of translating between programming languages differs from the challenge of translating natural languages in that programming languages are designed with a far more rigid s…