collaborators

6 papers

cs.CL2025

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs

Hyeonwoo Kim, Dahyun Kim, Jihoo Kim +3

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative…

cs.CL2025

Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models

Hyunbyung Park, Sukyung Lee, Gyoungjin Gim +3

To address the challenges associated with data processing at scale, we propose Dataverse, a unified open-source Extract-Transform-Load (ETL) pipeline for large language models (LLM…

cs.CL2024

Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models

Dahyun Kim, Sukyung Lee, Yungi Kim +2

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, a…

cs.CL2024

Evalverse: Unified and Accessible Library for Large Language Model Evaluation

Jihoo Kim, Wonho Song, Dahyun Kim +3

This paper introduces Evalverse, a novel library that streamlines the evaluation of Large Language Models (LLMs) by unifying disparate evaluation tools into a single, user-friendly…

cs.CL2024

sDPO: Don't Use Your Data All at Once

Dahyun Kim, Yungi Kim, Wonho Song +4

As development of large language models (LLM) progresses, aligning them with human preferences has become increasingly important. We propose stepwise DPO (sDPO), an extension of th…

cs.CL2024

1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models

Chanjun Park, Hyunsoo Ha, Jihoo Kim +4

In this paper, we propose the 1 Trillion Token Platform (1TT Platform), a novel framework designed to facilitate efficient data sharing with a transparent and equitable profit-shar…