1 citations · 1 across the 2 of their papers we have counts for
5 papers
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
Hoyoon Byun, Youngjun Choi, Taero Kim +2
Pre-Layer Normalization (Pre-LN) is the de facto choice for large language models (LLMs) and is crucial for stable pretraining and effective transfer learning. However, Pre-LN incu…
MIDUS: Memory-Infused Depth Up-Scaling
Taero Kim, Hoyoon Byun, Youngjun Choi +2
Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scaling (DUS) does so by duplicating Transfo…
LBC: Language-Based-Classifier for Out-Of-Variable Generalization
Kangjun Noh, Baekryun Seong, Hoyoon Byun +3
Large Language Models (LLMs) have great success in natural language processing tasks such as response generation. However, their use in tabular data has been limited due to their i…
Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments
Heeyoung Lee, Hoyoon Byun, Changdae Oh +2
Accessing machine learning models through remote APIs has been gaining prevalence following the recent trend of scaling up model parameters for increased performance. Even though t…
Flat Posterior Does Matter For Bayesian Model Averaging
Sungjun Lim, Jeyoon Yeom, Sooyon Kim +5
Bayesian neural networks (BNNs) estimate the posterior distribution of model parameters and utilize posterior samples for Bayesian Model Averaging (BMA) in prediction. However, des…