2 papers
cs.CL2025
A methodological analysis of prompt perturbations and their effect on attack success rates
Tiago Machado, Maysa Malfiza Garcia de Macedo, Rogerio Abreu de Paula +5
This work aims to investigate how different Large Language Models (LLMs) alignment methods affect the models' responses to prompt attacks. We selected open source models based on t…
cs.CL2025
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
Muneeza Azmat, Momin Abbas, Maysa Malfiza Garcia de Macedo +9
As Large Language Models (LLMs) become increasingly integrated into real-world applications, ensuring their outputs align with human values and safety standards has become critical…