collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes

Kyunghoon Bae, Eunbi Choi, Kibong Choi +37

This technical report introduces EXAONE 4.0, which integrates a Non-reasoning mode and a Reasoning mode to achieve both the excellent usability of EXAONE 3.5 and the advanced reaso…

cs.CL2024

LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models

Yungi Kim, Hyunsoo Ha, Seonghoon Yang +3

Creating high-quality, large-scale datasets for large language models (LLMs) often relies on resource-intensive, GPU-accelerated models for quality filtering, making the process ti…

cs.CL2024

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs

Hyeonwoo Kim, Dahyun Kim, Jihoo Kim +3

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative…

cs.CL2024

1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models

Chanjun Park, Hyunsoo Ha, Jihoo Kim +4

In this paper, we propose the 1 Trillion Token Platform (1TT Platform), a novel framework designed to facilitate efficient data sharing with a transparent and equitable profit-shar…

cs.CL2024

Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora

Yungi Kim, Hyunsoo Ha, Sukyung Lee +3

With the increasing demand for substantial amounts of high-quality data to train large language models (LLMs), efficiently filtering large web corpora has become a critical challen…