5 papers
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
Kai Xiong, Yanwei Huang, Rongjunchen Zhang +3
High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language models (LLMs). While recent data…
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
Kla Tantithamthavorn, Yaotian Zou, Andy Wong +8
Large Language Models (LLMs)-powered code review automation has the potential to transform code review workflows. Despite the advances of LLM-powered code review comment generation…
Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
Linfeng Zhang, Siheng Chen, Yuzhu Cai +46
AI agents are emerging as a practical way to run multi-step scientific workflows that interleave reasoning with tool use and verification, pointing to a shift from isolated AI-assi…
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
Tingjia Miao, Jiawen Dai, Jingkun Liu +23
Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain ex…
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
Danqing Zhang, Balaji Rama, Jingyi Ni +5
We introduce LiteWebAgent, an open-source suite for VLM-based web agent applications. Our framework addresses a critical gap in the web agent ecosystem with a production-ready solu…