6 papers
CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents
Taeyun Roh, Wonjune Jang, Junha Jung +1
Large language model agents heavily rely on external memory to support knowledge reuse and complex reasoning tasks. Yet most memory systems store experiences in a single global ret…
BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA
Taeyun Roh, Suhyeong Park, Eun-yeong Jo +5
Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs). However, because the M…
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
Yukyung Lee, Joonghoon Kim, Jaehee Kim +4
Existing LLM-as-a-Judge approaches for evaluating text generation suffer from rating inconsistencies, with low agreement and high rating variance across different evaluator models.…
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
Yukyung Lee, Soonwon Ka, Bokyung Son +2
Large Language Models (LLMs) have impacted the writing process, enhancing productivity by collaborating with humans in content creation platforms. However, generating high-quality,…
DSAI: Unbiased and Interpretable Latent Feature Extraction for Data-Centric AI
Hyowon Cho, Soonwon Ka, Daechul Park +3
Large language models (LLMs) often struggle to objectively identify latent characteristics in large datasets due to their reliance on pre-trained knowledge rather than actual data…
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
Hoonick Lee, Mogan Gim, Donghyeon Park +2
The advent of Large Language Models (LLMs) have shown promise in various creative domains, including culinary arts. However, many LLMs still struggle to deliver the desired level o…