1 citations · 1 across the 6 of their papers we have counts for
6 papers · 1 filter
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
Jiashu Yao, Heyan Huang, Zeming Liu +1
To overcome the sparse reward challenge in reinforcement learning (RL) for agents based on large language models (LLMs), we propose Mutual Information Self-Evaluation (MISE), an RL…
Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
Jiashu Yao, Heyan Huang, Shuang Zeng +6
Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on compl…
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
Silin Li, Yuhang Guo, Jiashu Yao +2
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, whi…
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks
Jiashu Yao, Heyan Huang, Zeming Liu +4
Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which…
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
Yuli Qiu, Jiashu Yao, Heyan Huang +1
Multi-step reasoning ability of large language models is crucial in tasks such as math and tool utilization. Current researches predominantly focus on enhancing model performance i…
FAME: Towards Factual Multi-Task Model Editing
Li Zeng, Yingyu Shan, Zeming Liu +2
Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks. Nevertheless, outdated knowledge or factual errors within…