activity
20142025
most citedContinual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities

7 citations · 34 across the 21 of their papers we have counts for

collaborators

16 papers

cs.CL20247 cited

Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities

Kazuki Fujii, Taishi Nakamura, Mengsay Loem +7

Cross-lingual continual pre-training of large language models (LLMs) initially trained on English corpus allows us to leverage the vast amount of English language resources and red…

cs.CL20245 cited

Building a Large Japanese Web Corpus for Large Language Models

Naoaki Okazaki, Kakeru Hattori, Hirai Shota +7

Open Japanese large language models (LLMs) have been trained on the Japanese portions of corpora such as CC-100, mC4, and OSCAR. However, these corpora were not created for the qua…

math.NA20231 cited

Computing the k-th Eigenvalue of Symmetric -Matrices

M. Ridwan Apriansyah, Rio Yokota

The numerical solution of eigenvalue problems is essential in various application areas of scientific and engineering domains. In many problem classes, the practical interest is on…

cs.PF2023

Cache Optimization and Performance Modeling of Batched, Small, and Rectangular Matrix Multiplication on Intel, AMD, and Fujitsu Processors

Sameer Deshmukh, Rio Yokota, George Bosilca

Factorization and multiplication of dense matrices and tensors are critical, yet extremely expensive pieces of the scientific toolbox. Careful use of low rank approximation can dra…

math.NA2023

distributed direct factorization of structured dense matrices using runtime systems

Sameer Deshmukh, Qinxiang Ma, Rio Yokota +1

Structured dense matrices result from boundary integral problems in electrostatics and geostatistics, and also Schur complements in sparse preconditioners such as multi-frontal met…

cs.CV20231 cited

SegRCDB: Semantic Segmentation via Formula-Driven Supervised Learning

Risa Shinoda, Ryo Hayamizu, Kodai Nakashima +3

Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks…