3 papers
cs.HC2026
Exploring Re-inforcement Learning via Human Feedback under User Heterogeneity
Sarvesh Shashidhar, Abhishek Mishra, Madhav Kotecha
Re-inforcement learning from human feedback (RLHF) has been effective in the task of AI alignment. However, one of the key assumptions of RLHF is that the annotators (referred to a…
cs.IR2026
FAIR-MATCH: A Multi-Objective Framework for Bias Mitigation in Reciprocal Dating Recommendations
Madhav Kotecha
Online dating platforms have fundamentally transformed the formation of romantic relationships, with millions of users worldwide relying on algorithmic matching systems to find com…
cs.LG2025
Subset Selection for Fine-Tuning: A Utility-Diversity Balanced Approach for Mathematical Domain Adaptation
Madhav Kotecha, Vijendra Kumar Vaishya, Smita Gautam +1
We propose a refined approach to efficiently fine-tune large language models (LLMs) on specific domains like the mathematical domain by employing a budgeted subset selection method…