2 papers
cs.AI2026
Aletheia tackles FirstProof autonomously
Tony Feng, Junehyuk Jung, Sang-hyun Kim +14
We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed t…
cs.LG2025
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
Allen Nie, Yi Su, Bo Chang +4
Despite their success in many domains, large language models (LLMs) remain under-studied in scenarios requiring optimal decision-making under uncertainty. This is crucial as many r…