9 papers
Detecting Out-of-Distribution Objects through Class-Conditioned Inpainting
Quang-Huy Nguyen, Jin Peng Zhou, Zhenzhen Liu +4
Recent object detectors have achieved impressive accuracy in identifying objects seen during training. However, real-world deployment often introduces novel and unexpected objects,…
: Provably Optimal Distributional RL for LLM Post-Training
Jin Peng Zhou, Kaiwen Wang, Jonathan Chang +5
Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inh…
Pre-training Limited Memory Language Models with Internal and External Knowledge
Linxi Zhao, Sofian Zalouk, Christian K. Belardi +7
Neural language models are black-boxes--both linguistic patterns and factual knowledge are distributed across billions of opaque parameters. This entangled encoding makes it diffic…
Learning to decode logical circuits
Yiqing Zhou, Chao Wan, Yichen Xu +3
With the development of quantum hardware bringing the error-corrected quantum circuits to the near future, the lack of an efficient polynomial-time decoding algorithms for logical…
INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations
Qian Meng, Jin Peng Zhou, Kilian Q. Weinberger +1
This paper presents INPROVF, an automatic framework that combines large language models (LLMs) and formal methods to speed up the repair process of high-level robot controllers. Pr…
Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
Qizhou Wang, Jin Peng Zhou, Zhanke Zhou +3
Large language models (LLMs) should undergo rigorous audits to identify potential risks, such as copyright and privacy infringements. Once these risks emerge, timely updates are cr…