activity
20242026
collaborators

13 papers

cs.CL2026

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

Jianing Wang, Xintao Wang, Aili Chen +7

Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to bui…

cs.AI2026

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Zehao Wang, Lanjun Wang, Shilong Jin +2

Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reason…

cs.CL2026

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Jinyi Han, Yuanjian Xu, Ying Liao +6

Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evalu…

cs.CL2026

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

Ying He, Zhouhong Gu, Zhecheng Hu +8

Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language…

cs.CL2026

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

Shixin Fang, Jiachen Wo, Wenjuan Qin +2

Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation li…

cs.LG2026

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

Zishang Jiang, Tingyun Li, Jinyi Han +7

Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face…