most citedRM-R1: Reward Modeling as Reasoning

2 citations · 4 across the 31 of their papers we have counts for

collaborators

34 papers

cs.CL2026

PolicyMem: Geometric Policy Memory for LLM Governance

Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao +3

As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two…

cs.CL2026

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

Haobo Xu, Sirui Chen, Yuanchen Bei +5

Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit…

cs.MA2026

One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles

Zhichen Zeng, Huiyuan Chen, Jingru Cheng +7

Specializing Large Language Models (LLMs) toward distinct abilities underpins successes ranging from personalized assistants to multi-agent systems (MAS). Single-agent paradigms re…

cs.CL2026

Beyond LLM-Based Reasoning: Lightweight GNNs for Agent Failure Attribution

Ting-Wei Li, Yuanchen Bei, Xiao Lin +1

Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task…

cs.CV2026

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

Qi Yu, Zhichen Zeng, Katherine Tieu +8

Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant effo…

cs.LG2026

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

Xuying Ning, Dongqi Fu, Tianxin Wei +13

Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions.…