5 papers
Understanding helpfulness and harmless tension in reward models
Eshaan Tanwar, Pepa Atanasova
Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal m…
Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends
Sabine Weber, Angelina Wang, Ankush Gupta +16
Natural language processing (NLP) technologies are rapidly reshaping how language is created, processed, and interpreted by humans. With current and potential applications in hirin…
Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing
Eshaan Tanwar, Gayatri Oke, Tanmoy Chakraborty
Bilingual lexical processing is shaped by the complex interplay of phonological, orthographic, and semantic features of two languages within an integrated mental lexicon. In humans…
Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
Eshaan Tanwar, Anwoy Chatterjee, Michael Saxon +3
Most multilingual question-answering benchmarks, while covering a diverse pool of languages, do not factor in regional diversity in the information they capture and tend to be West…
Understanding the Effects of Domain Finetuning on LLMs
Eshaan Tanwar, Deepak Nathani, William Yang Wang +1
Large Language Models (LLMs) fine-tuned for specific domains exhibit strong performance; however, the underlying mechanisms by which this fine-tuning reshapes their parametric spac…