7 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models
Boyan Han, Yiwei Wang, Yi Song +2
Diffusion large language models (dLLMs) offer bidirectional attention and parallel generation, enabling them to exploit global context and naturally support format-constrained task…
AppAgent-Claw: CLI Is All You Need for GUI Automation
Zhixue Song, Zhiheng Zhang, Yi Song +1
The OpenClaw platform provides a practical foundation for automation through its skill-oriented architecture, organizing external capabilities into lightweight, reusable components…
Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
Yuheng Yang, Wenjia Jiang, Yang Wang +3
The rapid progress of large language models (LLMs) has opened new opportunities for education. While learners can interact with academic papers through LLM-powered dialogue, limita…
SQLAgent: Learning to Explore Before Generating as a Data Engineer
Wenjia Jiang, Yiwei Wang, Boyan Han +2
Large Language Models have recently shown impressive capabilities in reasoning and code generation, making them promising tools for natural language interfaces to relational databa…
Learning to Be A Doctor: Searching for Effective Medical Agent Architectures
Yangyang Zhuang, Wenjia Jiang, Jiayu Zhang +3
Large Language Model (LLM)-based agents have demonstrated strong capabilities across a wide range of tasks, and their application in the medical domain holds particular promise due…