7 citations · 16 across the 26 of their papers we have counts for
10 papers · 1 filter
NARRA-Gym for Evaluating Interactive Narrative Agents
Yue Huang, Yuchen Ma, Jiayi Ye +14
Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limit…
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
Han Bao, Penghao Zhang, Yue Huang +9
Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability to comprehend and reason abou…
AI Alignment Breaks at the Edge
Han Bao, Yue Huang, Xiaoda Wang +5
General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. The problem is not only that mode…
ProbeLLM: Automating Principled Diagnosis of LLM Failures
Yue Huang, Zhengzhe Jiang, Yuchen Ma +8
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has…
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
Haomin Zhuang, Yujun Zhou, Taicheng Guo +4
Reinforcement Learning has demonstrated substantial improvements in the reasoning abilities of Large Language Models (LLMs), exhibiting significant applicability across various dom…
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
Yue Huang, Zhengzhe Jiang, Xiaonan Luo +12
Empowering large language models (LLMs) with chemical intelligence remains a challenge due to the scarcity of high-quality, domain-specific instruction-response datasets and the mi…