collaborators

6 papers

cs.LG2026

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

Yiran Shen, Yu Xia, Jonathan Chang +1

Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer wha…

cs.LG2026

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization

Rohan Surana, Xintong Li, Sheldon Yu +7

Multi-negative preference optimization under the Plackett--Luce (PL) model extends Direct Preference Optimization (DPO) by leveraging comparative signals across one preferred and m…

cs.SD2025

Preference-Based Learning in Audio Applications: A Systematic Analysis

Aaron Broukhim, Yiran Shen, Prithviraj Ammanabrolu +1

Despite the parallel challenges that audio and text domains face in evaluating generative model outputs, preference learning remains remarkably underexplored in audio applications.…

cs.LG2025

In-context Ranking Preference Optimization

Junda Wu, Rohan Surana, Zhouhang Xie +6

Recent developments in Direct Preference Optimization (DPO) allow large language models (LLMs) to function as implicit ranking models by maximizing the margin between preferred and…

cs.CL2025

SAND: Boosting LLM Agents with Self-Taught Action Deliberation

Yu Xia, Yiran Shen, Junda Wu +5

Large Language Model (LLM) agents are commonly tuned with supervised finetuning on ReAct-style expert trajectories or preference optimization over pairwise rollouts. Most of these…

cs.CL2025

A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

Zhouhang Xie, Junda Wu, Yiran Shen +9

Personalized preference alignment for large language models (LLMs), the process of tailoring LLMs to individual users' preferences, is an emerging research direction spanning the a…