activity
20242026
most citedCausal Evaluation of Language Models

3 citations · 7 across the 22 of their papers we have counts for

collaborators
Showing 2026Show all

9 papers · 1 filter

cs.LG2026

Safin-1: Safety from Within through Memory-Native State Evolution

Ming Zhang, Kaisen Yang, Shu Yu +15

Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic proper…

cs.LG2026

CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery

Bo Peng, Kaiwen Wu, Sirui Chen +3

Causal discovery from observational data remains challenging due to the fundamental limitations of purely statistical methods, such as statistical distinguishability within equival…

cs.CL2026

Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals

Sirui Chen, Lei Xu, Yuying Zhao +6

Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewar…

cs.CL2026

Code as Agent Harness

Xuying Ning, Katherine Tieu, Dongqi Fu +39

Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineerin…

cs.RO2026

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

Harold Haodong Chen, Sirui Chen, Yingjie Xu +2

The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video gener…

cs.CL2026

EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

Ting-Wei Li, Sirui Chen, Jiaru Zou +4

Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires iteratively improving the model…