5 papers
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
Jefferson Hernandez, Jaywon Koo, Zilin Xiao +2
Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group normalization induces a diverge…
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
Jaywon Koo, Jefferson Hernandez, Ruozhen He +3
We introduce HypoExplore, an agentic framework that formulates neural architecture discovery for visual recognition as a hypothesis-driven scientific inquiry. Given a human-specifi…
A Mathematical Theory of Agency and Intelligence
Wael Hafez, Chenan Wei, Rodrigo Pena +2
To operate reliably under changing conditions, complex systems require feedback on how effectively they use resources, not just whether objectives are met. Current AI systems proce…
Play to Generalize: Learning to Reason Through Game Play
Yunfei Xie, Yinsong Ma, Shiyi Lan +3
Developing reasoning capabilities in multimodal large language models (MLLMs) remains challenging. Motivated by literature suggesting that gameplay promotes transferable reasoning…
GenEx: Generating an Explorable World
Taiming Lu, Tianmin Shu, Junfei Xiao +8
Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step to…