6 papers
CoFrGeNet: Continued Fraction Architectures for Language Generation
Amit Dhurandhar, Vijil Chenthamarakshan, Dennis Wei +3
Transformers are arguably the preferred architecture for language generation. In this paper, inspired by continued fractions, we introduce a new function class for generative model…
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
Zahra Ashktorab, Michael Desmond, Qian Pan +7
Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
Paying Alignment Tax with Contrastive Learning
Buse Sibel Korkmaz, Rahul Nair, Elizabeth M. Daly +1
Current debiasing approaches often result a degradation in model capabilities such as factual accuracy and knowledge retention. Through systematic evaluation across multiple benchm…
Foundation Models at Work: Fine-Tuning for Fairness in Algorithmic Hiring
Buse Sibel Korkmaz, Rahul Nair, Elizabeth M. Daly +3
Foundation models require fine-tuning to ensure their generative outputs align with intended results for specific tasks. Automating this fine-tuning process is challenging, as it t…
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Nico Wagner, Michael Desmond, Rahul Nair +6
LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty…