11 papers
Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks
Boxiu Li, Zimo Wen, Yijia Fan +24
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a…
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
Linbin Tang, Jingyan You, Zilin Kang +8
Recent formal reasoning systems have reached IMO-level performance, yet they leave a fragmented landscape: algebra and number theory are handled in Lean, while geometry still relie…
STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
Jiaao Wu, Xian Zhang, Hanzhang Liu +3
Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, long-horizon reasoning, existin…
Draft-and-Prune: Improving the Reliability of Auto-formalization for Logical Reasoning
Zhiyu Ni, Zheng Liang, Liangcheng Song +4
Auto-formalization (AF) translates natural-language reasoning problems into solver-executable programs, enabling symbolic solvers to perform sound logical deduction. In practice, h…
PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
Yingfan Deng, Anhao Zhou, Yuan Yuan +3
Multi-Agent Systems have recently emerged as a promising paradigm for collaborative reasoning and solving complex tasks. However, the design of collaborative learning algorithms in…
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
Yunchi Lu, Youshan Miao, Cheng Tan +4
Training large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed traini…