works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

q-bio.QM2026

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Hyunjin Seo, Hyeon Hwang, Gyubok Lee +7

The paper introduces TheBioCollection, a 52.6‑billion‑token unified corpus that aggregates diverse biological resources for pre‑training large language models, and shows that train…

cs.LG2026

Predicting LLM Reasoning Performance with Small Proxy Model

Woosung Koh, Juyoung Suk, Sungjun Han +2

Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets before scaling up. However, this approach be…

cs.LG2026

Generative Visual Code Mobile World Models

Woosung Koh, Sungjun Han, Segyu Lee +2

Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and inference-time. However, current approaches…

q-bio.QM2026

VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

Hyunjin Seo, Hongjoon Ahn, Jimin Park +16

Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly sh…

cs.CL2025

Trillion 7B Technical Report

Sungjun Han, Juyoung Suk, Suyeong An +5

We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient a…

cs.LG2025

Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning

Seyeon Kim, Joonhun Lee, Namhoon Cho +2

Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residua…