activity
20162024
most citedEvaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

51 citations · 97 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2024

The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models

Ryosuke Takahashi, Go Kamoda, Benjamin Heinzerling +2

Language models (LMs) encode world knowledge in their internal parameters through training. However, LMs may learn personal and confidential information from the training data, lea…

cs.CL2024

J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema

Kosuke Matsuzaki, Masaya Taniguchi, Kentaro Inui +1

We introduce a Japanese Morphology dataset, J-UniMorph, developed based on the UniMorph feature schema. This dataset addresses the unique and rich verb forms characteristic of the…

cs.CL2023

Test-time Augmentation for Factual Probing

Go Kamoda, Benjamin Heinzerling, Keisuke Sakaguchi +1

Factual probing is a method that uses prompts to test if a language model "knows" certain world knowledge facts. A problem in factual probing is that small changes to the prompt ca…

cs.CL202351 cited

Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

Jungo Kasai, Yuhei Kasai, Keisuke Sakaguchi +2

As large language models (LLMs) gain popularity among speakers of diverse languages, we believe that it is crucial to benchmark them to better understand model behaviors, failures,…

cs.CL20231 cited

Causal schema induction for knowledge discovery

Michael Regan, Jena D. Hwang, Keisuke Sakaguchi +1

Making sense of familiar yet new situations typically involves making generalizations about causal schemas, stories that help humans reason about event sequences. Reasoning about e…

cs.CL202327 cited

Analyzing the Performance of GPT-3.5 and GPT-4 in Grammatical Error Correction

Steven Coyne, Keisuke Sakaguchi, Diana Galvan-Sosa +2

GPT-3 and GPT-4 models are powerful, achieving high performance on a variety of Natural Language Processing tasks. However, there is a relative lack of detailed published analysis…