6 papers
Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
Heloisa Candello, Muneeza Azmat, Uma Sushmitha Gunturi +7
With the rapid uptake of generative AI, investigating human perceptions of generated responses has become crucial. A major challenge is their `aptitude' for hallucinating and gener…
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
Paulo Cavalin, Cassia Sanctos, Marcelo Grave +2
We introduce \textsc{CAT}, a framework designed to evaluate and visualize the \emph{interplay} of \emph{accuracy} and \emph{response consistency} of Large Language Models (LLMs) un…
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
Paulo Cavalin, Cassia Sanctos, Marcelo Grave +2
In this work we present the Consistency-Rebalanced Accuracy (CoRA) metric, improving the reliability of Large Language Model (LLM) scores computed on multiple choice (MC) benchmark…
A methodological analysis of prompt perturbations and their effect on attack success rates
Tiago Machado, Maysa Malfiza Garcia de Macedo, Rogerio Abreu de Paula +5
This work aims to investigate how different Large Language Models (LLMs) alignment methods affect the models' responses to prompt attacks. We selected open source models based on t…
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Claudio Pinhanez, Paulo Cavalin, Cassia Sanctos +2
This work explores the consistency of small LLMs (2B-8B parameters) in answering multiple times the same question. We present a study on known, open-source LLMs responding to 10 re…
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
Muneeza Azmat, Momin Abbas, Maysa Malfiza Garcia de Macedo +9
As Large Language Models (LLMs) become increasingly integrated into real-world applications, ensuring their outputs align with human values and safety standards has become critical…