3 papers
cs.SD2026
ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
Jisu Jeon, Seungyeon Jwa, Joosung Lee +6
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holisti…
cs.CL2025
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
Hwiyeol Jo, Joosung Lee, Jaehone Lee +3
Evaluating generative models, such as large language models (LLMs), commonly involves question-answering tasks where the final answer is selected based on probability of answer cho…
cs.CL2025
Enhancing Hallucination Detection via Future Context
Joosung Lee, Cheonbok Park, Hwiyeol Jo +3
Large Language Models (LLMs) are widely used to generate plausible text on online platforms, without revealing the generation process. As users increasingly encounter such black-bo…