activity
20242026
collaborators

10 papers

cs.AI2026

Towards Direct Evaluation of Harness Optimizers via Priority Ranking

Kai Tzu-iunn Ong, Minseok Kang, Dongwook Choi +9

Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate op…

cs.AI2026

PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints

Minjun Park, Donghyun Kim, Hyeonjong Ju +5

We are entering an era in which individuals and organizations increasingly deploy dedicated AI agents that interact and collaborate with other agents. However, the dynamics of mult…

cs.CL2026

EXAONE 4.5 Technical Report

Eunbi Choi, Kibong Choi, Sehyun Chun +55

This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integrating a dedicated visual enc…

cs.AI2026

Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers

Wooseok Seo, Seungju Han, Jaehun Jung +6

Fact verification is essential for ensuring the reliability of LLM applications. In this study, we evaluate 12 pre-trained LLMs and one specialized fact-verifier, including frontie…

cs.CL2025

Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games

Seungwon Lim, Seungbeen Lee, Dongjun Min +1

Artificial agents are increasingly central to complex interactions and decision-making tasks, yet aligning their behaviors with desired human values remains an open challenge. In t…

cs.AI2025

VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms

Seungwon Lim, Sungwoong Kim, Jihwan Yu +3

Escape rooms present a unique cognitive challenge that demands exploration-driven planning: with the sole instruction to 'escape the room', players must actively search their envir…