works on

From the 1 of 56 linked papers with an AI index.

activity
20242026
most citedRepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository

1 citations · 2 across the 19 of their papers we have counts for

collaborators
Showing cs.AIShow all

17 papers · 1 filter

cs.AI2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Sungho Park, Wonjoong Kim, Rongyuan Tan +10

LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses…

cs.AI2026

TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search

Zhuofan Shi, Mingzhe Ma, Lu Wang +8

Deep search requires agents to answer complex questions through multi-step web search, browsing, evidence comparison, and synthesis. A central challenge is deciding how to search w…

cs.AI20261 cited

StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis

Jiayi Mao, Liqun Li, Yanjie Gao +9

Effective incident management in large-scale IT systems relies on troubleshooting guides (TSGs), but their manual execution is slow and error-prone. While recent advances in LLMs o…

cs.AI2026

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

Rongyuan Tan, Jue Zhang, Zhuozhao Li +3

Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or toy settings, leaving their beh…

cs.AI2026

WebXSkill: Skill Learning for Autonomous Web Agents

Zhaoyang Wang, Qianhui Wu, Xuchao Zhang +12

Autonomous web agents powered by large language models (LLMs) have shown promise in completing complex browser tasks, yet they still struggle with long-horizon workflows. A key bot…

cs.AI2026

DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems

Ming Ma, Jue Zhang, Fangkai Yang +4

Large language model (LLM)-based multi-agent systems are challenging to debug because failures often arise from long, branching interaction traces. The prevailing practice is to le…