7 papers
Beyond Expectations: Learning with Stochastic Dominance Made Practical
Shicong Cen, Jincheng Mei, Hanjun Dai +3
Stochastic dominance serves as a general framework for modeling a broad spectrum of decision preferences under uncertainty, with risk aversion as one notable example, as it natural…
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
Changhao Li, Yuchen Zhuang, Rushi Qiang +4
Despite the impressive generative abilities of black-box large language models (LLMs), their inherent opacity hinders further advancements in capabilities such as reasoning, planni…
AmorLIP: Efficient Language-Image Pretraining via Amortization
Haotian Sun, Yitong Li, Yuchen Zhuang +3
Contrastive Language-Image Pretraining (CLIP) has demonstrated strong zero-shot performance across diverse downstream text-image tasks. Existing CLIP methods typically optimize a c…
Towards Better Instruction Following Retrieval Models
Yuchen Zhuang, Aaron Trinh, Rushi Qiang +4
Modern information retrieval (IR) models, trained exclusively on standard <query, passage> pairs, struggle to effectively interpret and follow explicit user instructions. We introd…
Faster WIND: Accelerating Iterative Best-of- Distillation for LLM Alignment
Tong Yang, Jincheng Mei, Hanjun Dai +5
Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algo…
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Shicong Cen, Jincheng Mei, Katayoon Goshvadi +6
Reinforcement learning from human feedback (RLHF) has demonstrated great promise in aligning large language models (LLMs) with human preference. Depending on the availability of pr…