activity
20242026
collaborators

7 papers

cs.LG2026

Beyond Expectations: Learning with Stochastic Dominance Made Practical

Shicong Cen, Jincheng Mei, Hanjun Dai +3

Stochastic dominance serves as a general framework for modeling a broad spectrum of decision preferences under uncertainty, with risk aversion as one notable example, as it natural…

cs.LG2025

Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs

Changhao Li, Yuchen Zhuang, Rushi Qiang +4

Despite the impressive generative abilities of black-box large language models (LLMs), their inherent opacity hinders further advancements in capabilities such as reasoning, planni…

cs.LG2025

AmorLIP: Efficient Language-Image Pretraining via Amortization

Haotian Sun, Yitong Li, Yuchen Zhuang +3

Contrastive Language-Image Pretraining (CLIP) has demonstrated strong zero-shot performance across diverse downstream text-image tasks. Existing CLIP methods typically optimize a c…

cs.CL2025

Towards Better Instruction Following Retrieval Models

Yuchen Zhuang, Aaron Trinh, Rushi Qiang +4

Modern information retrieval (IR) models, trained exclusively on standard <query, passage> pairs, struggle to effectively interpret and follow explicit user instructions. We introd…

cs.LG2025

Faster WIND: Accelerating Iterative Best-of- Distillation for LLM Alignment

Tong Yang, Jincheng Mei, Hanjun Dai +5

Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algo…

cs.LG2025

Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Shicong Cen, Jincheng Mei, Katayoon Goshvadi +6

Reinforcement learning from human feedback (RLHF) has demonstrated great promise in aligning large language models (LLMs) with human preference. Depending on the availability of pr…