Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Aletheia tackles FirstProof autonomously
Tony Feng, Junehyuk Jung, Sang-hyun Kim +14
We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed t…
cs.AI2026
Understanding the Role of Training Data in Test-Time Scaling
Adel Javanmard, Baharan Mirzasoleiman, Vahab Mirrokni
Test-time scaling improves the reasoning capabilities of large language models (LLMs) by allocating extra compute to generate longer Chains-of-Thoughts (CoTs). This enables models…