3 papers
cs.LG2026
The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
Eric Yeats, Brendan Kennedy, Loc Truong +5
The hidden states of large language models (LLMs) are known to capture rich information relating to model knowledge and behavior that can be hard to extract from examination of inp…
cs.LG2024
Fine-Grained Uncertainty Quantification via Collisions
Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon
We propose a new and intuitive metric for aleatoric uncertainty quantification (UQ), the prevalence of class collisions defined as the same input being observed in different classe…
cs.LG2024
Trustworthy Actionable Perturbations
Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon
Counterfactuals, or modified inputs that lead to a different outcome, are an important tool for understanding the logic used by machine learning classifiers and how to change an un…