13 citations · 15 across the 3 of their papers we have counts for
2 papers
cs.LG2025
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee +5
Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we in…
cs.CL2019★ 13 cited
Why Do Masked Neural Language Models Still Need Common Sense Knowledge?
Sunjae Kwon, Cheongwoong Kang, Jiyeon Han +1
Currently, contextualized word representations are learned by intricate neural network models, such as masked neural language models (MNLMs). The new representations significantly…