17 citations · 74 across the 22 of their papers we have counts for
27 papers
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
Shirui Chen, Jiantao Jiao, Lillian J. Ratliff +1
Masked diffusion language models (MDLMs) offer the potential for parallel token generation, but most open-source MDLMs decode fewer than 5 tokens per model forward pass even with s…
Efficient Prompt Caching via Embedding Similarity
Hanlin Zhu, Banghua Zhu, Jiantao Jiao
Large language models (LLMs) have achieved huge success in numerous natural language process (NLP) tasks. However, it faces the challenge of significant resource consumption during…
Generative AI Security: Challenges and Countermeasures
Banghua Zhu, Norman Mu, Jiantao Jiao +1
Generative AI's expanding footprint across numerous industries has led to both excitement and increased scrutiny. This paper delves into the unique security challenges posed by Gen…
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Banghua Zhu, Michael I. Jordan, Jiantao Jiao
Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique that aligns language models closely with human-centric values. The initial phase of RLHF involves learning…
Towards Optimal Statistical Watermarking
Baihe Huang, Hanlin Zhu, Banghua Zhu +4
We study statistical watermarking by formulating it as a hypothesis testing problem, a general framework which subsumes all previous statistical watermarking methods. Key to our fo…
Guided Online Distillation: Promoting Safe Reinforcement Learning by Offline Demonstration
Jinning Li, Xinyi Liu, Banghua Zhu +4
Safe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints. When learning from scratch, safe RL agents tend to be overly co…