3 citations · 8 across the 13 of their papers we have counts for
6 papers · 1 filter
Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization
Travis Zhang, Christian Belardi, Justin Lovelace +4
Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. While much work has focused on e…
Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
Qizhou Wang, Jin Peng Zhou, Zhanke Zhou +3
Large language models (LLMs) should undergo rigorous audits to identify potential risks, such as copyright and privacy infringements. Once these risks emerge, timely updates are cr…
Graders should cheat: privileged information enables expert-level automated evaluations
Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding +3
Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with i…
: Provably Optimal Distributional RL for LLM Post-Training
Jin Peng Zhou, Kaiwen Wang, Jonathan Chang +5
Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inh…
Detecting Out-of-Distribution Objects through Class-Conditioned Inpainting
Quang-Huy Nguyen, Jin Peng Zhou, Zhenzhen Liu +4
Recent object detectors have achieved impressive accuracy in identifying objects seen during training. However, real-world deployment often introduces novel and unexpected objects,…
Does Label Differential Privacy Prevent Label Inference Attacks?
Ruihan Wu, Jin Peng Zhou, Kilian Q. Weinberger +1
Label differential privacy (label-DP) is a popular framework for training private ML models on datasets with public features and sensitive private labels. Despite its rigorous priv…