10 papers
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Jeonghwan Choi, Taewon Yun, Minjeong Ban +3
Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide…
SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations
Taewon Yun, Hyeonseong Park, Jeonghwan Choi +3
Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rel…
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
Taewon Yun, Jisu Shin, Jeonghwan Choi +2
Distilling large reasoning models is essential for making Long-CoT reasoning practical, as full-scale inference remains computationally prohibitive. Existing curation-based approac…
What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents
Jeonghwan Choi, Jibin Hwang, Gyeonghun Sun +4
Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion through buyer-seller interaction to purc…
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
Jiaqi Deng, Yuho Lee, Nicole Hee-Yeon Kim +5
We introduce HAMLET, a holistic and automated framework for evaluating the long-context comprehension of large language models (LLMs). HAMLET structures source texts into a three-l…
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
Hyangsuk Min, Yuho Lee, Minjeong Ban +6
Evaluation frameworks for text summarization have evolved in terms of both domain coverage and metrics. However, existing benchmarks still lack domain-specific assessment criteria,…