6 citations · 20 across the 30 of their papers we have counts for
9 papers · 1 filter
MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
Tal Kadosh, Niranjan Hasabnis, Vy A. Vo +10
With easier access to powerful compute resources, there is a growing trend in AI for software development to develop large language models (LLMs) to address a variety of programmin…
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
Anaelia Ovalle, Ninareh Mehrabi, Palash Goyal +6
Gender-inclusive NLP research has documented the harmful limitations of gender binary-centric large language models (LLM), such as the inability to correctly use gender-diverse Eng…
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
Stephen Mayhew, Terra Blevins, Shuheng Liu +10
We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-qual…
Analyzing Cognitive Plausibility of Subword Tokenization
Lisa Beinborn, Yuval Pinter
Subword tokenization has become the de-facto standard for tokenization, although comparative evaluations of subword vocabulary quality across languages are scarce. Existing evaluat…
Emptying the Ocean with a Spoon: Should We Edit Models?
Yuval Pinter, Michael Elhadad
We call into question the recently popularized method of direct model editing as a means of correcting factual errors in LLM generations. We contrast model editing with three simil…
Scope is all you need: Transforming LLMs for HPC Code
Tal Kadosh, Niranjan Hasabnis, Vy A. Vo +9
With easier access to powerful compute resources, there is a growing trend in the field of AI for software development to develop larger and larger language models (LLMs) to addres…