activity
20192024
most citedGitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical Errors

17 citations · 25 across the 9 of their papers we have counts for

collaborators

11 papers

cs.CL2024

Project MOSLA: Recording Every Moment of Second Language Acquisition

Masato Hagiwara, Joshua Tanner

Second language acquisition (SLA) is a complex and dynamic process. Many SLA studies that have attempted to record and analyze this process have typically focused on a single modal…

cs.SD20221 cited

AVES: Animal Vocalization Encoder based on Self-Supervision

Masato Hagiwara

The lack of annotated training data in bioacoustics hinders the use of large-scale neural network models trained in a supervised way. In order to leverage a large amount of unannot…

cs.SD20221 cited

BEANS: The Benchmark of Animal Sounds

Masato Hagiwara, Benjamin Hoffman, Jen-Yu Liu +3

The use of machine learning (ML) based techniques has become increasingly popular in the field of bioacoustics over the last years. Fundamental requirements for the successful appl…

cs.SD20221 cited

Modeling Animal Vocalizations through Synthesizers

Masato Hagiwara, Maddie Cusimano, Jen-Yu Liu

Modeling real-world sound is a fundamental problem in the creative use of machine learning and many other fields, including human speech processing and bioacoustics. Transformer-ba…

cs.CL20223 cited

Towards Automated Document Revision: Grammatical Error Correction, Fluency Edits, and Beyond

Masato Mita, Keisuke Sakaguchi, Masato Hagiwara +3

Natural language processing technology has rapidly improved automated grammatical error correction tasks, and the community begins to explore document-level revision as one of the…

cs.CL2021

Semi-Supervised Joint Estimation of Word and Document Readability

Yoshinari Fujinuma, Masato Hagiwara

Readability or difficulty estimation of words and documents has been investigated independently in the literature, often assuming the existence of extensive annotated resources for…