3 papers
cs.AI2025
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
Roshita Bhonsle, Rishav Dutta, Sneha Vavilapalli +8
The increasing adoption of foundation models as agents across diverse domains necessitates a robust evaluation framework. Current methods, such as LLM-as-a-Judge, focus only on fin…
cs.CL2024
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
Anmol Mekala, Vineeth Dorna, Shreya Dubey +5
Machine unlearning aims to efficiently eliminate the influence of specific training data, known as the forget set, from the model. However, existing unlearning methods for Large La…
cs.CL2024
Does Prompt Formatting Have Any Impact on LLM Performance?
Jia He, Mukund Rungta, David Koleczek +3
In the realm of Large Language Models (LLMs), prompt optimization is crucial for model performance. Although previous research has explored aspects like rephrasing prompt contexts,…