activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits

Qianshu Cai, Yonggang Zhang, Jun Nie +6

Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in response to task feedback while keeping the underlying language model…

cs.AI2026

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Zhiqin Yang, Jingwen Fu, Yuhan Liu +16

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code,…

cs.AI2026

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

Chunyang Jiang, Pingping Zhang, Yuzhi Zhao +9

Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimod…

cs.AI2026

MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems

Qianshu Cai, Yonggang Zhang, Xianzhang Jia +5

Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a…

cs.AI2026

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

Zhiqin Yang, Yonggang Zhang, Wei Xue +3

Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implem…