From the 1 of 6 linked papers with an AI index.
6 papers
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
Hyunjin Seo, Hyeon Hwang, Gyubok Lee +7
The paper introduces TheBioCollection, a 52.6‑billion‑token unified corpus that aggregates diverse biological resources for pre‑training large language models, and shows that train…
Predicting LLM Reasoning Performance with Small Proxy Model
Woosung Koh, Juyoung Suk, Sungjun Han +2
Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets before scaling up. However, this approach be…
Generative Visual Code Mobile World Models
Woosung Koh, Sungjun Han, Segyu Lee +2
Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and inference-time. However, current approaches…
VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
Hyunjin Seo, Hongjoon Ahn, Jimin Park +16
Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly sh…
Trillion 7B Technical Report
Sungjun Han, Juyoung Suk, Suyeong An +5
We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient a…
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
Seyeon Kim, Joonhun Lee, Namhoon Cho +2
Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residua…