4 papers
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
Mokshit Surana, Archit Rathod, Akshaj Satishkumar
Large Language Models (LLMs) trained on web-scale corpora inherently absorb toxic patterns from their training data. This leads to toxic degeneration where even innocuous prompts c…
Fair and Calibrated Toxicity Detection with Robust Training and Abstention
Mokshit Surana
Fairness in toxicity classification involves three integrated axes: ranking, calibration, and abstention. Training-time interventions and post-hoc safety mechanisms cannot be evalu…
A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
Pranav Pushkar Mishra, Kranti Prakash Yeole, Ramyashree Keshavamurthy +2
In enterprise settings, efficiently retrieving relevant information from large and complex knowledge bases is essential for operational productivity and informed decision-making. T…
Examining the Implications of Deepfakes for Election Integrity
Hriday Ranka, Mokshit Surana, Neel Kothari +7
It is becoming cheaper to launch disinformation operations at scale using AI-generated content, in particular 'deepfake' technology. We have observed instances of deepfakes in poli…