3 papers
cs.LG2026
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
Ziru Chen, Dongdong Chen, Ruinan Jin +3
Recently, there have been significant research interests in training large language models (LLMs) with reinforcement learning (RL) on real-world tasks, such as multi-turn code gene…
cs.AI2025
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…
astro-ph.IM2025
Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)
Lucas Carrit Delgado Pinheiro, Ziru Chen, Bruno Caixeta Piazza +4
While task-specific demonstrations show early success in applying large language models (LLMs) to automate some astronomical research tasks, they only provide incomplete views of a…