Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation
Ashima Sood, Bryan Gardiner, Joan Condell
Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture help…
cs.CL2026
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
Karamvir Singh Batra, Prathamjyot Singh, Ashima Sood +2
At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of…