activity
20242026
most citedAgentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Yang Liu, Bin Chong, Wenkai Yang +9

Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to…

cs.AI2026

MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

Jiajun Dong, Yutao Hu, Fengrui Fan +3

Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-rele…

cs.AI2026

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

Ziliang Zhao, Zenan Xu, Shuting Wang +7

Planning is a fundamental capability for large language models (LLMs) because such complex tasks require models to coordinate goals, constraints, resources, and long-term consequen…

cs.AI2026

Towards Security-Auditable LLM Agents: A Unified Graph Representation

Chaofan Li, Lyuye Zhang, Jintao Zhai +9

LLM-based agentic systems are rapidly evolving to perform complex autonomous tasks through dynamic tool invocation, stateful memory management, and multi-agent collaboration. Howev…

cs.AI2026

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

Yuhui Wang, Zhixiong Yang, Ming Zhang +10

In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by la…

cs.AI2026

Steering LLMs via Scalable Interactive Oversight

Enyu Zhou, Zhiheng Xi, Long Ma +9

As Large Language Models increasingly automate complex, long-horizon tasks such as \emph{vibe coding}, a supervision gap has emerged. While models excel at execution, users often s…