activity
20202026
most citedLLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs

5 citations · 6 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda +6

Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a m…

cs.CL2026

Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition

Shiho Matta, Yin Jou Huang, Fei Cheng +3

We introduce Dango, a 1.8B-parameter large language model designed for controlled studies of L1-to-L2 (Japanese-to-English) transfer in second language acquisition (SLA). While pre…

cs.CL2024★ 1 cited

Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis

Shiho Matta, Yin Jou Huang, Fei Cheng +2

Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may no…

cs.CL2024★ 5 cited

LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs

LLM-jp, :, Akiko Aizawa +80

This paper introduces LLM-jp, a cross-organizational project for the research and development of Japanese large language models (LLMs). LLM-jp aims to develop open-source and stron…

cs.CL2024

RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State

Takashi Kodama, Hirokazu Kiyomaru, Yin Jou Huang +1

Humans pay careful attention to the interlocutor's internal state in dialogues. For example, in recommendation dialogues, we make recommendations while estimating the seeker's inte…

cs.CL2023

MultiTool-CoT: GPT-3 Can Use Multiple External Tools with Chain of Thought Prompting

Tatsuro Inaba, Hirokazu Kiyomaru, Fei Cheng +1

Large language models (LLMs) have achieved impressive performance on various reasoning tasks. To further improve the performance, we propose MultiTool-CoT, a novel framework that l…