activity
20242026
collaborators

5 papers

cs.CL2026

Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models

Avni Mittal, Shanu Kumar, Sandipan Dandapat +1

We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are missing. This problem is comm…

cs.CR2025

SAGE: A Generic Framework for LLM Safety Evaluation

Madhur Jindal, Hari Shrawgi, Parag Agrawal +1

As Large Language Models are rapidly deployed across diverse applications from healthcare to financial advice, safety evaluation struggles to keep pace. Current benchmarks focus on…

cs.CY2025

LLM Safety for Children

Prasanjit Rath, Hari Shrawgi, Parag Agrawal +1

This paper analyzes the safety of Large Language Models (LLMs) in interactions with children below age of 18 years. Despite the transformative applications of LLMs in various aspec…

cs.CL2024

Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation

Shanu Kumar, Gauri Kholkar, Saish Mendke +3

With the growth of social media and large language models, content moderation has become crucial. Many existing datasets lack adequate representation of different groups, resulting…

cs.CL2024

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection

Shanu Kumar, Saish Mendke, Karody Lubna Abdul Rahman +3

Chain-of-thought (CoT) prompting has significantly enhanced the capability of large language models (LLMs) by structuring their reasoning processes. However, existing methods face…