6 citations · 18 across the 13 of their papers we have counts for
13 papers
Is Micro Domain-Adaptive Pre-Training Effective for Real-World Operations? Multi-Step Evaluation Reveals Potential and Bottlenecks
Masaya Tsunokake, Yuta Koreeda, Terufumi Morishita +3
When applying LLMs to real-world enterprise operations, LLMs need to handle proprietary knowledge in small domains of specific operations (). A previous stu…
Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio +1
Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the reliance on costly specialized target l…
A Nested Watermark for Large Language Models
Koichi Nagatsuka, Terufumi Morishita, Yasuhiro Sogawa
The rapid advancement of large language models (LLMs) has raised concerns regarding their potential misuse, particularly in generating fake news and misinformation. To address thes…
Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic Corpus
Terufumi Morishita, Gaku Morio, Atsuki Yamaguchi +1
Large language models (LLMs) are capable of solving a wide range of tasks, yet they have struggled with reasoning. To address this, we propose $\textbf{Additional Logic Training (A…
Adapting Chat Language Models Using Only Target Unlabeled Language Data
Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio +1
Vocabulary expansion (VE) is the de-facto approach to language adaptation of large language models (LLMs) by adding new tokens and continuing pre-training on target data. While thi…
appjsonify: An Academic Paper PDF-to-JSON Conversion Toolkit
Atsuki Yamaguchi, Terufumi Morishita
We present appjsonify, a Python-based PDF-to-JSON conversion toolkit for academic papers. It parses a PDF file using several visual-based document layout analysis models and rule-b…