6 papers
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Jeonghwan Choi, Taewon Yun, Minjeong Ban +3
Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide…
Solar Open 2 Technical Report
Sungrae Park, Sanghoon Kim, Gyoungjin Gim +50
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent tra…
What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents
Jeonghwan Choi, Jibin Hwang, Gyeonghun Sun +4
Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion through buyer-seller interaction to purc…
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
Minjeong Ban, Jeonghwan Choi, Hyangsuk Min +4
Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies…
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
Jiaqi Deng, Yuho Lee, Nicole Hee-Yeon Kim +5
We introduce HAMLET, a holistic and automated framework for evaluating the long-context comprehension of large language models (LLMs). HAMLET structures source texts into a three-l…
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
Hyangsuk Min, Yuho Lee, Minjeong Ban +6
Evaluation frameworks for text summarization have evolved in terms of both domain coverage and metrics. However, existing benchmarks still lack domain-specific assessment criteria,…