4 papers
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
Hoyoon Byun, Youngjun Choi, Taero Kim +2
Pre-Layer Normalization (Pre-LN) is the de facto choice for large language models (LLMs) and is crucial for stable pretraining and effective transfer learning. However, Pre-LN incu…
MIDUS: Memory-Infused Depth Up-Scaling
Taero Kim, Hoyoon Byun, Youngjun Choi +2
Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scaling (DUS) does so by duplicating Transfo…
Flat Posterior Does Matter For Bayesian Model Averaging
Sungjun Lim, Jeyoon Yeom, Sooyon Kim +5
Bayesian neural networks (BNNs) estimate the posterior distribution of model parameters and utilize posterior samples for Bayesian Model Averaging (BMA) in prediction. However, des…
Graph Perceiver IO: A General Architecture for Graph Structured Data
Seyun Bae, Hoyoon Byun, Changdae Oh +2
Multimodal machine learning has been widely studied for the development of general intelligence. Recently, the Perceiver and Perceiver IO, show competitive results for diverse data…