3 papers
cs.AI2026
Inference-Time Nash Alignment
Hadi Hosseini, Debmalya Mandal, Duohan Zhang
Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are…
cs.LG2024
Bandit Learning in Matching Markets: Utilitarian and Rawlsian Perspectives
Hadi Hosseini, Duohan Zhang
Two-sided matching markets have demonstrated significant impact in many real-world applications, including school choice, medical residency placement, electric vehicle charging, ri…
cs.GT2024
Putting Gale & Shapley to Work: Guaranteeing Stability Through Learning
Hadi Hosseini, Sanjukta Roy, Duohan Zhang
Two-sided matching markets describe a large class of problems wherein participants from one side of the market must be matched to those from the other side according to their prefe…