1 citations · 1 across the 1 of their papers we have counts for
3 papers
Language Modeling With Factorization Memory
Lee Xiong, Maksim Tkachenko, Johanes Effendi +1
We propose Factorization Memory, an efficient recurrent neural network (RNN) architecture that achieves performance comparable to Transformer models on short-context language model…
Revisiting Transformers with Insights from Image Filtering and Boosting
Laziz U. Abdullaev, Maksim Tkachenko, Tan M. Nguyen
The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpre…
RakutenAI-7B: Extending Large Language Models for Japanese
Rakuten Group, Aaron Levine, Connie Huang +27
We introduce RakutenAI-7B, a suite of Japanese-oriented large language models that achieve the best performance on the Japanese LM Harness benchmarks among the open 7B models. Alon…