activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang

Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexam…

cs.CL2025

Towards Transparent AI: A Survey on Explainable Language Models

Avash Palikhe, Zichong Wang, Zhipeng Yin +4

Language Models (LMs) have significantly advanced natural language processing and enabled remarkable progress across diverse domains, yet their black-box nature raises critical con…

cs.CL2025

Towards Transparent AI: A Survey on Explainable Large Language Models

Avash Palikhe, Zhenyu Yu, Zichong Wang +1

Large Language Models (LLMs) have played a pivotal role in advancing Artificial Intelligence (AI). However, despite their achievements, LLMs often struggle to explain their decisio…

cs.CL2025

Datasets for Fairness in Language Models: An In-Depth Survey

Jiale Zhang, Zichong Wang, Avash Palikhe +2

Despite the growing reliance on fairness benchmarks to evaluate language models, the datasets that underpin these benchmarks remain critically underexamined. This survey addresses…

cs.CL2024

Fairness in Large Language Models in Three Hours

Thang Doan Viet, Zichong Wang, Minh Nhat Nguyen +1

Large Language Models (LLMs) have demonstrated remarkable success across various domains but often lack fairness considerations, potentially leading to discriminatory outcomes agai…

cs.CL2024

Fairness Definitions in Language Models Explained

Zhipeng Yin, Zichong Wang, Avash Palikhe +1

Language Models (LMs) have demonstrated exceptional performance across various Natural Language Processing (NLP) tasks. Despite these advancements, LMs can inherit and amplify soci…