10 citations · 16 across the 19 of their papers we have counts for
7 papers · 1 filter
Learning the Error Patterns of Language Models
Jinwoo Kim, Taylor Berg-KirkPatrick, Loris D'Antoni
When generating outputs for domains with specific validity constraints (e.g., a program should compile), LLMs often fail in a small number of focused ways: for example, by using Py…
Manifold-Guided Attention Steering
Ian Li, Kapilesh Guruprasad, Raunak Sengupta +3
Large language models frequently produce errors in reasoning tasks despite possessing the underlying knowledge required for correct reasoning. One possible approach to improve reas…
Continuous Diffusion Models Can Obey Formal Syntax
Jinwoo Kim, Taylor Berg-Kirkpatrick, Loris D'Antoni
Diffusion language models offer a promising alternative to autoregressive models due to their global, non-causal generation process, but their continuous latent dynamics make discr…
Verified Training for Counterfactual Explanation Robustness under Data Shift
Anna P. Meyer, Yuhao Zhang, Aws Albarghouthi +1
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by describing what changes to an input are necessary to change its prediction to a desired…
Certifying Robustness to Programmable Data Bias in Decision Trees
Anna P. Meyer, Aws Albarghouthi, Loris D'Antoni
Datasets can be biased due to societal inequities, human biases, under-representation of minorities, etc. Our goal is to certify that models produced by a learning algorithm are po…
Certified Robustness to Programmable Transformations in LSTMs
Yuhao Zhang, Aws Albarghouthi, Loris D'Antoni
Deep neural networks for natural language processing are fragile in the face of adversarial examples -- small input perturbations, like synonym substitution or word duplication, wh…