146 citations · 166 across the 8 of their papers we have counts for
9 papers
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditio…
A Mechanistic View of Authority Hierarchy in LLM Sycophancy
Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answe…
Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models
Priyanka Mary Mammen, Emil Joswin, Shankar Venkitachalam
Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However, the influence of endorsement s…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…
Detecting Natural Language Biases with Prompt-based Learning
Md Abdul Aowal, Maliha T Islam, Priyanka Mary Mammen +1
In this project, we want to explore the newly emerging field of prompt engineering and apply it to the downstream task of detecting LM biases. More concretely, we explore how to de…