7 papers
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure that the retained state influences subsequen…
Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadratic accumulation of independent scalars ($\…
Z-Plane Neural Networks: Bounded Geometric Activation Replaces ReLU and LayerNorm
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Modern deep neural networks rely on Euclidean scalar activations (e.g., ReLU) and global normalization techniques (e.g., LayerNorm) to prevent gradient instability in deep architec…
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
For over a decade, explicit memory architectures like the Neural Turing Machine have remained theoretically appealing yet practically intractable for language modeling due to catas…
LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
Jong Myoung Kim, Young-Jun Lee, Ho-Jin Choi +1
While Large Language Models have gained attention, many service developers still rely on embedding-based models due to practical constraints. In such cases, the quality of fine-tun…
PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment
Jong Myoung Kim, Young-Jun_Lee, Ho-Jin Choi +1
Transfer learning leverages the abundance of English data to address the scarcity of resources in modeling non-English languages, such as Korean. In this study, we explore the pote…