collaborators

6 papers

cs.LG2026

Bradley-Terry Policy Optimization for Generative Preference Modeling

Shengyu Feng, Yun He, Shuang Ma +12

Reinforcement learning (RL) has recently proven effective at scaling chain-of-thought (CoT) reasoning in large language models for tasks with verifiable answers. However, extending…

cs.AI2026

Generalized Parallel Scaling with Interdependent Generations

Harry Dong, David Brandfonbrener, Eryk Helenowski +5

Parallel LLM inference scaling involves sampling a set of responses for a single input prompt. However, these parallel responses tend to be generated independently from e…

cs.AI2025

Reinforcement Learning from User Feedback

Eric Han, Jun Chen, Karthik Abinav Sankararaman +8

As large language models (LLMs) are increasingly deployed in diverse user facing applications, aligning them with real user preferences becomes essential. Existing methods like Rei…

cs.CL2025

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation

Chengwei Qin, Wenxuan Zhou, Karthik Abinav Sankararaman +10

Hallucination, the generation of factually incorrect information, remains a significant challenge for large language models (LLMs), especially in open-domain long-form generation.…

cs.LG2025

Preference Optimization with Multi-Sample Comparisons

Chaoqi Wang, Zhuokai Zhao, Chen Zhu +8

Recent advancements in generative models, particularly large language models (LLMs) and diffusion models, have been driven by extensive pretraining on large datasets followed by po…

cs.AI2025

Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

Zishun Yu, Tengyu Xu, Di Jin +9

Solving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as…