32 citations · 34 across the 4 of their papers we have counts for
4 papers
Morphological evaluation of subwords vocabulary used by BETO language model
Óscar García-Sierra, Ana Fernández-Pampillón Cesteros, Miguel Ortega-Martín
Subword tokenization algorithms used by Large Language Models are significantly more efficient and can independently build the necessary vocabulary of words and subwords without hu…
Building another Spanish dictionary, this time with GPT-4
Miguel Ortega-Martín, Óscar García-Sierra, Alfonso Ardoiz +8
We present the "Spanish Built Factual Freectianary 2.0" (Spanish-BFF-2) as the second iteration of an AI-generated Spanish dictionary. Previously, we developed the inaugural versio…
Spanish Built Factual Freectianary (Spanish-BFF): the first AI-generated free dictionary
Miguel Ortega-Martín, Óscar García-Sierra, Alfonso Ardoiz +3
Dictionaries are one of the oldest and most used linguistic resources. Building them is a complex task that, to the best of our knowledge, has yet to be explored with generative La…
Linguistic ambiguity analysis in ChatGPT
Miguel Ortega-Martín, Óscar García-Sierra, Alfonso Ardoiz +3
Linguistic ambiguity is and has always been one of the main challenges in Natural Language Processing (NLP) systems. Modern Transformer architectures like BERT, T5 or more recently…