6 papers
Triplet loss based embeddings for forensic speaker identification in Spanish
Emmanuel Maqueda, Javier Alvarez-Jimenez, Carlos Mena +1
With the advent of digital technology, it is more common that committed crimes or legal disputes involve some form of speech recording where the identity of a speaker is questioned…
Hacia los Comités de Ética en Inteligencia Artificial
Sofía Trejo, Ivan Meza, Fernanda López-Escobedo
The goal of Artificial Intelligence based systems is to take decisions that have an effect in their environment and impact society. This points out to the necessity of mechanism th…
Topic Discovery in Massive Text Corpora Based on Min-Hashing
Gibran Fuentes-Pineda, Ivan Vladimir Meza-Ruiz
The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to mas…
Lost in Translation: Analysis of Information Loss During Machine Translation Between Polysynthetic and Fusional Languages
Manuel Mager, Elisabeth Mager, Alfonso Medina-Urrea +2
Machine translation from polysynthetic to fusional languages is a challenging task, which gets further complicated by the limited amount of parallel text available. Thus, translati…
Challenges of language technologies for the indigenous languages of the Americas
Manuel Mager, Ximena Gutierrez-Vasques, Gerardo Sierra +1
Indigenous languages of the American continent are highly diverse. However, they have received little attention from the technological perspective. In this paper, we review the res…
Fortification of Neural Morphological Segmentation Models for Polysynthetic Minimal-Resource Languages
Katharina Kann, Manuel Mager, Ivan Meza-Ruiz +1
Morphological segmentation for polysynthetic languages is challenging, because a word may consist of many individual morphemes and training data can be extremely scarce. Since neur…