2 papers
cs.CL2025
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
Kohei Tsuji, Tatsuya Hiraoka, Yuchang Cheng +2
This paper investigates how LLMs encode inputs with typos. We hypothesize that specific neurons and attention heads recognize typos and fix them internally using local and global c…
cs.CL2025
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
Kohei Tsuji, Tatsuya Hiraoka, Yuchang Cheng +1
NLP datasets may still contain annotation errors, even when they are manually annotated. Researchers have attempted to develop methods to automatically reduce the adverse effect of…