5 papers
Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability
Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi +1
Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previous works examined the effects…
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
Krishnapriya Vishnubhotla, Sowmya Vajjala, Akriti Vij +1
We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that Large Language Models are u…
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
Isar Nejadgholi, Masoud Kianpour, Krishnapriya Vishnubhotla +1
Tremendous efforts have been put into evaluating the inclusivity and effectiveness of AI systems across cultures. However, the cultural capabilities considered in much of the liter…
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
Xi Yu Huang, Krishnapriya Vishnubhotla, Frank Rudzicz
The improved generative capabilities of large language models have made them a powerful tool for creative writing and storytelling. It is therefore important to quantitatively unde…