activity
20222026
most citedA Unified Evaluation of Textual Backdoor Learning: Frameworks and Benchmarks

24 citations · 54 across the 30 of their papers we have counts for

collaborators
Showing 2026 · cs.AIShow all

9 papers · 2 filters

cs.AI2026

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

Yinghao Chen, Zixi Chen, Bingxiang He +7

Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the…

cs.AI2026

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Yizhe Chi, Wenyi Li, Deyao Hong +7

Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the t…

cs.AI2026

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Yuhao Zhan, Bingxiang He, Zecong Tang +1

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery afte…

cs.AI2026

A-Evolve-Training: Autonomous Post-Training of a 30B Model

Zhan Shi, Bing He, Yisi Sang +2

Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous sys…

cs.AI2026

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10

Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains under…

cs.AI2026

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

Yizhe Chi, Deyao Hong, Dapeng Jiang +18

Current LLM agent benchmarks, which predominantly focus on binary pass/fail tasks such as code generation or search-based question answering, often neglect the value of real-world…