7 papers
MolDA: Molecular Understanding and Generation via Large Language Diffusion Model
Seohyeon Shin, HanJun Choi, Jun-Hyung Park +2
Large Language Models (LLMs) have significantly advanced molecular discovery, but existing multimodal molecular architectures fundamentally rely on autoregressive (AR) backbones. T…
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
Nayeon Kim, Eojin Jeon, Jun-Hyung Park +1
In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a…
Incorporating Domain Knowledge into Materials Tokenization
Yerim Oh, Jun-Hyung Park, Junho Kim +2
While language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language proces…
C2A: Client-Customized Adaptation for Parameter-Efficient Federated Learning
Yeachan Kim, Junho Kim, Wing-Lam Mok +2
Despite the versatility of pre-trained language models (PLMs) across domains, their large memory footprints pose significant challenges in federated learning (FL), where the traini…
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
Junho Kim, Yeachan Kim, Jun-Hyung Park +3
We introduce a novel continued pre-training method, MELT (MatEriaLs-aware continued pre-Training), specifically designed to efficiently adapt the pre-trained language models (PLMs)…
Zero-shot Commonsense Reasoning over Machine Imagination
Hyuntae Park, Yeachan Kim, Jun-Hyung Park +1
Recent approaches to zero-shot commonsense reasoning have enabled Pre-trained Language Models (PLMs) to learn a broad range of commonsense knowledge without being tailored to speci…