6 papers
Capacity-Dependent Effects of Data Selection for Reasoning
Cuong Dang, Hoang Anh Just, Ruoxi Jia
In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelih…
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
Aashish Dhawan, Christopher Driggers-Ellis, Dzmitry Kasinets +2
We present the University of Florida Gators submission to the AmericasNLP 2026 shared task on cultural image captioning for Indigenous languages. Our two-stage pipeline generates a…
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
Christopher Driggers-Ellis, Nachiketh Tibrewal, Rohit Bogulla +4
A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exis…
RISE: Interactive Visual Diagnosis of Fairness in Machine Learning Models
Ray Chen, Christan Grant
Evaluating fairness under domain shift is challenging because scalar metrics often obscure exactly where and how disparities arise. We introduce \textit{RISE} (Residual Inspection…
Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
Aashish Dhawan, Christopher Driggers-Ellis, Christan Grant +1
Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for…
MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
Christopher Driggers-Ellis, Detravious Brinkley, Ray Chen +3
Multi30k is frequently cited in the multimodal machine translation (MMT) literature, offering parallel text data for training and fine-tuning deep learning models. However, it is l…