4 papers
KAIO: A Collection of More Challenging Korean Questions
Nahyun Lee, Guijin Son, Hyunwoo Ko +1
With the advancement of mid/post-training techniques, LLMs are pushing their boundaries at an accelerated pace. Legacy benchmarks saturate quickly (e.g., broad suites like MMLU ove…
Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
Kyubeen Han, Junseo Jang, Hongjin Kim +2
Instruction-tuning enhances the ability of large language models (LLMs) to follow user instructions more accurately, improving usability while reducing harmful outputs. However, th…
EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
Kyunghoon Bae, Eunbi Choi, Kibong Choi +37
This technical report introduces EXAONE 4.0, which integrates a Non-reasoning mode and a Reasoning mode to achieve both the excellent usability of EXAONE 3.5 and the advanced reaso…
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour +9
Faithfulness evaluators based on large language models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries. We propose an appro…