activity
20242026
most citedHomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation

Jiashu Yao, Heyan Huang, Zeming Liu +1

To overcome the sparse reward challenge in reinforcement learning (RL) for agents based on large language models (LLMs), we propose Mutual Information Self-Evaluation (MISE), an RL…

cs.CL2025

Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement

Jiashu Yao, Heyan Huang, Shuang Zeng +6

Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on compl…

cs.CL20251 cited

HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices

Silin Li, Yuhang Guo, Jiashu Yao +2

Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, whi…

cs.CL2024

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks

Jiashu Yao, Heyan Huang, Zeming Liu +4

Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which…

cs.CL2024

Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation

Yuli Qiu, Jiashu Yao, Heyan Huang +1

Multi-step reasoning ability of large language models is crucial in tasks such as math and tool utilization. Current researches predominantly focus on enhancing model performance i…

cs.CL2024

FAME: Towards Factual Multi-Task Model Editing

Li Zeng, Yingyu Shan, Zeming Liu +2

Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks. Nevertheless, outdated knowledge or factual errors within…