3 papers
cs.CL2025
A methodological analysis of prompt perturbations and their effect on attack success rates
Tiago Machado, Maysa Malfiza Garcia de Macedo, Rogerio Abreu de Paula +5
This work aims to investigate how different Large Language Models (LLMs) alignment methods affect the models' responses to prompt attacks. We selected open source models based on t…
cs.CL2025
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
Muneeza Azmat, Momin Abbas, Maysa Malfiza Garcia de Macedo +9
As Large Language Models (LLMs) become increasingly integrated into real-world applications, ensuring their outputs align with human values and safety standards has become critical…
cs.CL2025
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
Prashanth Vijayaraghavan, Soroush Vosoughi, Lamogha Chiazor +4
Recent advancements in large language models (LLMs) have revolutionized natural language processing (NLP) and expanded their applications across diverse domains. However, despite t…