most citedAutonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

Qing Zong, Jiayu Liu, Junhao Shen +9

Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedb…

cs.CL2026

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Jiayu Liu, Cheng Qian, Zhenhailong Wang +10

Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are progressively disclosed through…

cs.CL2026

PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning

Bingxuan Li, Jeonghwan Kim, Cheng Qian +4

Overlapping calendar invitations force busy professionals to repeatedly decide which meetings to attend, reschedule, or decline. We refer to this preference-driven decision process…

cs.CL2026

Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum

Gaotang Li, Ruizhong Qiu, Xiusi Chen +2

Supervised fine-tuning (SFT) is the standard approach for post-training large language models (LLMs), yet it often shows limited generalization. We trace this limitation to its def…

cs.CL2025

Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning

Qi He, Cheng Qian, Xiusi Chen +3

Claim verification with large language models (LLMs) has recently attracted growing attention, due to their strong reasoning capabilities and transparent verification processes com…

cs.CL2025

EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

Cheng Qian, Peixuan Han, Qinyu Luo +9

Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative ada…