3 papers
stat.ML2026
Adaptive Nucleus Truncation for Long-Form Reasoning
Ousmane Amadou Dia
Sampling plays an important role in long-form language-model reasoning. Over thousands of decoding steps, small changes in the candidate token set can compound into different reaso…
stat.ML2026
Variational Proximal Policy Optimization
Ousmane Amadou Dia
Reinforcement Learning from Human Feedback via Proximal Policy Optimization often suffers from policy mode collapse, brittle exploration loops, and distribution drift. This paper i…
cs.CL2026
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Yang Xu, Xuanming Zhang, Samuel Yeh +4
Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most eval…