collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Solar Open 2 Technical Report

Sungrae Park, Sanghoon Kim, Gyoungjin Gim +50

We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent tra…

cs.CL2026

Solar Open Technical Report

Sungrae Park, Sanghoon Kim, Jungho Cho +34

We introduce Solar Open, a 102B-parameter bilingual Mixture-of-Experts language model for underserved languages. Solar Open demonstrates a systematic methodology for building compe…

cs.CL2024

LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models

Yungi Kim, Hyunsoo Ha, Seonghoon Yang +3

Creating high-quality, large-scale datasets for large language models (LLMs) often relies on resource-intensive, GPU-accelerated models for quality filtering, making the process ti…

cs.CL2024

Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models

Dahyun Kim, Sukyung Lee, Yungi Kim +2

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, a…

cs.CL2024

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs

Hyeonwoo Kim, Dahyun Kim, Jihoo Kim +3

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative…

cs.CL2024

1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models

Chanjun Park, Hyunsoo Ha, Jihoo Kim +4

In this paper, we propose the 1 Trillion Token Platform (1TT Platform), a novel framework designed to facilitate efficient data sharing with a transparent and equitable profit-shar…