4 papers · 1 filter
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
Dongryeol Lee, Yerin Hwang, Yongil Kim +2
In line with the principle of honesty, there has been a growing effort to train large language models (LLMs) to generate outputs containing epistemic markers. However, evaluation i…
Return of EM: Entity-driven Answer Set Expansion for QA Evaluation
Dongryeol Lee, Minwoo Lee, Kyungmin Min +2
Recently, directly using large language models (LLMs) has been shown to be the most reliable method to evaluate QA models. However, it suffers from limited interpretability, high c…
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
Minbeom Kim, Hwanhee Lee, Joonsuk Park +2
As the integration of large language models into daily life is on the rise, there is a clear gap in benchmarks for advising on subjective and personal dilemmas. To address this, we…
LifeTox: Unveiling Implicit Toxicity in Life Advice
Minbeom Kim, Jahyun Koo, Hwanhee Lee +3
As large language models become increasingly integrated into daily life, detecting implicit toxicity across diverse contexts is crucial. To this end, we introduce LifeTox, a datase…