Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models
Wei Yu, Suxing Liu, Minjie Yu +4
Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We p…
cs.AI2026
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
Wei Yu, Suxing Liu, Minjie Yu +4
Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of sim…