activity
20242026
collaborators

9 papers

cs.CL2026

CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models

Yukyung Lee, Yumeng Shen, Jinhyeong Park +2

Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approaches typically lack alignment…

cs.AI2026

MolDA: Molecular Understanding and Generation via Large Language Diffusion Model

Seohyeon Shin, HanJun Choi, Jun-Hyung Park +2

Large Language Models (LLMs) have significantly advanced molecular discovery, but existing multimodal molecular architectures fundamentally rely on autoregressive (AR) backbones. T…

cs.CL2025

Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning

Nayeon Kim, Eojin Jeon, Jun-Hyung Park +1

In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a…

cs.CL2025

Incorporating Domain Knowledge into Materials Tokenization

Yerim Oh, Jun-Hyung Park, Junho Kim +2

While language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language proces…

cs.LG2024

C2A: Client-Customized Adaptation for Parameter-Efficient Federated Learning

Yeachan Kim, Junho Kim, Wing-Lam Mok +2

Despite the versatility of pre-trained language models (PLMs) across domains, their large memory footprints pose significant challenges in federated learning (FL), where the traini…

cs.CL2024

MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science

Junho Kim, Yeachan Kim, Jun-Hyung Park +3

We introduce a novel continued pre-training method, MELT (MatEriaLs-aware continued pre-Training), specifically designed to efficiently adapt the pre-trained language models (PLMs)…