5 papers · 1 filter
The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
Eric Yeats, Brendan Kennedy, Loc Truong +5
The hidden states of large language models (LLMs) are known to capture rich information relating to model knowledge and behavior that can be hard to extract from examination of inp…
What do Geometric Hallucination Detection Metrics Actually Measure?
Eric Yeats, John Buckheit, Sarah Scullen +9
Hallucination remains a barrier to deploying generative models in high-consequence applications. This is especially true in cases where external ground truth is not readily availab…
A Connection Between Score Matching and Local Intrinsic Dimension
Eric Yeats, Aaron Jacobson, Darryl Hannan +4
The local intrinsic dimension (LID) of data is a fundamental quantity in signal processing and learning theory, but quantifying the LID of high-dimensional, complex data has been a…
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
Eric Yeats, Darryl Hannan, Henry Kvinge +2
Machine unlearning (MU) is a promising cost-effective method to cleanse undesired information (generated concepts, biases, or patterns) from foundational diffusion models. While MU…
NashAE: Disentangling Representations through Adversarial Covariance Minimization
Eric Yeats, Frank Liu, David Womble +1
We present a self-supervised method to disentangle factors of variation in high-dimensional data that does not rely on prior knowledge of the underlying variation profile (e.g., no…