5 papers · 1 filter
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Hyangsuk Min, Hwanjun Song
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as str…
Solar Open 2 Technical Report
Sungrae Park, Sanghoon Kim, Gyoungjin Gim +50
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent tra…
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
Minjeong Ban, Jeonghwan Choi, Hyangsuk Min +4
Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies…
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
Hwanjun Song, Jeonghwan Choi, Minseok Kim
Retrieval-augmented generation (RAG) enhances LLMs with external knowledge, yet generation remains vulnerable to retrieval-induced noise and uncertain placement of relevant chunks,…
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
Jiaqi Deng, Yuho Lee, Nicole Hee-Yeon Kim +5
We introduce HAMLET, a holistic and automated framework for evaluating the long-context comprehension of large language models (LLMs). HAMLET structures source texts into a three-l…