2 papers
cs.CL2025
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
Austin Xu, Srijan Bansal, Yifei Ming +2
The large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-…
cs.CL2023
Few-shot Unified Question Answering: Tuning Models or Prompts?
Srijan Bansal, Semih Yavuz, Bo Pang +2
Question-answering (QA) tasks often investigate specific question types, knowledge domains, or reasoning skills, leading to specialized models catering to specific categories of QA…