3 papers
cs.LG2026
LLMs Should Express Uncertainty Explicitly
Junyu Guo, Shangding Gu, Ming Jin +2
Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…
cs.LG2026
StyleBench: Evaluating thinking styles in Large Language Models
Junyu Guo, Shangding Gu, Ming Jin +2
Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and control constraints. When additional reaso…
cs.LG2025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
Junyu Guo, Zhi Zheng, Donghao Ying +4
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common…