works on

From the 1 of 59 linked papers with an AI index.

activity
20242026
most citedRepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository

1 citations · 1 across the 18 of their papers we have counts for

collaborators
Showing cs.AIShow all

15 papers · 1 filter

cs.AI2026

TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search

Zhuofan Shi, Mingzhe Ma, Lu Wang +8

Deep search requires agents to answer complex questions through multi-step web search, browsing, evidence comparison, and synthesis. A central challenge is deciding how to search w…

cs.AI2026

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

Rongyuan Tan, Jue Zhang, Zhuozhao Li +3

Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or toy settings, leaving their beh…

cs.AI2026

DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems

Ming Ma, Jue Zhang, Fangkai Yang +4

Large language model (LLM)-based multi-agent systems are challenging to debug because failures often arise from long, branching interaction traces. The prevailing practice is to le…

cs.AI2025

GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents

Jian Mu, Chaoyun Zhang, Chiming Ni +14

We introduce GUI-360, a large-scale, comprehensive dataset and benchmark suite designed to advance computer-using agents (CUAs). CUAs present unique challenges and is const…

cs.AI2025

From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models

Jue Zhang, Qingwei Lin, Saravan Rajmohan +1

Large Reasoning Models (LRMs) generate explicit reasoning traces alongside final answers, yet the extent to which these traces influence answer generation remains unclear. In this…

cs.AI2025

Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation

Kaikai An, Fangkai Yang, Liqun Li +10

Recent advances in retrieval-augmented generation (RAG) have substantially improved question-answering systems, particularly for factoid '5Ws' questions. However, significant chall…