4 papers
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Yutong He, Daibo Li, Guohong Li +15
ReasFlow is an autonomous multi‑agent system that leverages large language models to perform rigorous mathematical reasoning, retrieve relevant knowledge, and generate complete res…
No Time Like the Present: Agentic Test-Time Training for LLM Agents
Yanbo Wang, Jinhua Hao, Yuze Shi +2
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time traini…
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
Xinyu Che, Junqi Xiong, Yunfei Ge +9
Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and i…
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama +1
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trai…