2 papers
cs.CL2026
Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation
Ayush Gupta, Hima Varshini Surisetty, Sreevidya Bollineni +5
Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has trul…
cs.LG2025
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
Tuhina Tripathi, Manya Wadhwa, Greg Durrett +1
Large Language Models (LLMs) are widely used as proxies for human labelers in both training (Reinforcement Learning from AI Feedback) and large-scale response evaluation (LLM-as-a-…