4 papers · 1 filter
Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan
Alexander Chulzhanov, Soeren Eberhardt, Arjun Mukherjee
Neural machine translation for digitally low-resource Indigenous languages is often hindered by extreme data scarcity, prompting reliance on extractive web-scraping. To ensure data…
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
Navid Ayoobi, Marcus I Armstrong, Arjun Mukherjee
Large language models (LLMs) reason over discrete token ID sequences, yet modern subword tokenizers routinely produce non-unique encodings: multiple token ID sequences can detokeni…
Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee +2
The local news landscape, a vital source of reliable information for 28 million Americans, faces a growing threat from Pink Slime Journalism, a low-quality, auto-generated articles…
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
Navid Ayoobi, Sadat Shahriar, Arjun Mukherjee
We present a novel evaluation paradigm for AI text detectors that prioritizes real-world and equitable assessment. Current approaches predominantly report conventional metrics like…