17 papers
Disentangling Geometry, Performance, and Training in Language Models
Atharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian +1
Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estima…
Token Rankings are Unforgeable Language Model Signatures
Matthew Finlayson, Andreas Grivas, Xiang Ren +1
Language model parameters are known to impose unique (to each model) geometric constraints on their logit outputs, which serves as a signature that identifies the model, but also l…
Side-by-side Comparison Amplifies Dialect Bias in Language Models
Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel +5
Language models (LMs) can exhibit biases based on variations in their dialects, even in the absence of a dialect label, a behavior known as covert dialect bias. In this work, we qu…
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46
This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
Yue Huang, Chujie Gao, Siyuan Wu +63
Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
Keyu He, Tejas Srinivasan, Brihi Joshi +3
When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM predictions with natural langua…