2 papers
cs.AI2026
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13
Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusabl…
cs.CL2025
Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training
Xinyuan Zhou, Yi Lei, Xiaoyu Zhou +7
Large Language Models (LLMs) have shown significant promise in automated theorem proving, yet progress is often constrained by the scarcity of diverse and high-quality formal langu…