Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
Hoyoon Byun, Youngjun Choi, Taero Kim +2
Pre-Layer Normalization (Pre-LN) is the de facto choice for large language models (LLMs) and is crucial for stable pretraining and effective transfer learning. However, Pre-LN incu…
cs.CL2024
LBC: Language-Based-Classifier for Out-Of-Variable Generalization
Kangjun Noh, Baekryun Seong, Hoyoon Byun +3
Large Language Models (LLMs) have great success in natural language processing tasks such as response generation. However, their use in tabular data has been limited due to their i…