collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Easper: An Accessible ASR Pipeline for Language Documentation

Aso Mahmudi, Ting Dang, Ekaterina Vylomova +1

Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists o…

cs.CL2026

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

David Setiawan, Temuulen Khishigsuren, Milind Agarwal +3

Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, th…

cs.CL2026

CommonMorph: Participatory Morphological Documentation Platform

Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan +3

Collecting and annotating morphological data present significant challenges, requiring linguistic expertise, methodological rigour, and substantial resources. These barriers are pa…

cs.CL2024

Can a Neural Model Guide Fieldwork? A Case Study on Morphological Data Collection

Aso Mahmudi, Borja Herce, Demian Inostroza Amestica +3

Linguistic fieldwork is an important component in language documentation and preservation. However, it is a long, exhaustive, and time-consuming process. This paper presents a nove…

cs.CL2024

Low-Resource Machine Translation through Retrieval-Augmented LLM Prompting: A Study on the Mambai Language

Raphaël Merx, Aso Mahmudi, Katrina Langford +2

This study explores the use of large language models (LLMs) for translating English into Mambai, a low-resource Austronesian language spoken in Timor-Leste, with approximately 200,…