3 citations · 3 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
Maheed H. Ahmed, Mahsa Ghasemi
Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents. However, in most scenarios, the mod…
cs.LG2025
Reinforcement Learning from Multi-level and Episodic Human Feedback
Muhammad Qasim Elahi, Somtochukwu Oguchienti, Maheed H. Ahmed +1
Designing an effective reward function has long been a challenge in reinforcement learning, particularly for complex tasks in unstructured environments. To address this, various le…