2 papers
cs.CL2026
Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users
Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari +2
To align a Large Language Model (LLM), most existing methods collect explicit human feedback and train a reward model to predict the human preference based on the response text. Th…
cs.CL2025
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
Aryan Sajith, Krishna Chaitanya Rao Kathala
This study investigates the relative impact of training data quality versus quantity on the performance of small language models (SLMs), utilizing the TinyStories dataset for empir…