4 citations · 8 across the 8 of their papers we have counts for
9 papers
PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
Wesley Hanwen Deng, Sunnie S. Y. Kim, Akshita Jha +4
Recent developments in AI governance and safety research have called for red-teaming methods that can effectively surface potential risks posed by AI models. Many of these calls ha…
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
Sanchit Kabra, Akshita Jha, Chandan K. Reddy
Recent advances in large-scale generative language models have shown that reasoning capabilities can significantly improve model performance across a variety of tasks. However, the…
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
Akshita Jha, Sanchit Kabra, Chandan K. Reddy
Recent studies have shown that generative language models often reflect and amplify societal biases in their outputs. However, these studies frequently conflate observed biases wit…
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton +5
Recent studies have shown that Text-to-Image (T2I) model generations can reflect social stereotypes present in the real world. However, existing approaches for evaluating stereotyp…
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models
Akshita Jha, Aida Davani, Chandan K. Reddy +3
Stereotype benchmark datasets are crucial to detect and mitigate social stereotypes about groups of people in NLP models. However, existing datasets are limited in size and coverag…
Transformer-based Models for Long-Form Document Matching: Challenges and Empirical Analysis
Akshita Jha, Adithya Samavedhi, Vineeth Rakesh +2
Recent advances in the area of long document matching have primarily focused on using transformer-based models for long document encoding and matching. There are two primary challe…