#llm agents

topicllm agents

49 papers · 1 filter

cs.CL2026

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Albert Gong, Kyuseong Choi, Abhineet Agarwal +5

The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…

cs.AI2026

MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck

Dongyi Liu, Haixing He, Xiaobao Wu +1

The paper introduces MIND, a lightweight framework that uses an intent‑aware information bottleneck to detect and filter poisoned memory in large language model agents, reducing at…

cs.AR2026

ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents

Stef Cuyckens, Mihaela Jivanescu, Jun Yin +2

The paper presents ARES, a framework that adaptively controls the reasoning effort of large language model agents when optimizing RTL designs for power, performance, and area, whil…

cs.CL2026

ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory

Yongye Su, Wujiang Xu, Chaoji Zuo +1

ChronoMem adds a semantic version‑control layer to large language model agents, allowing them to snapshot, browse, and roll back their long‑term memory using natural‑language reque…

cs.AI2026

Baikal: Structured Search for Deep Research over Data Lakes

Dhruv Agarwal, Rishitha Guttapalle Mohan, Aarti Kumari +5

Baikal is a framework that clusters heterogeneous tables and passages into semantic regions and uses adaptive, budgeted search policies to guide an LLM agent in generating subquest…

cs.MA2026

Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling

Zuyuan Zhang, Hanqing Yang, Carlee Joe-Wong +1

The paper proposes iCORE, a unified representation that combines a cooperation graph, an obligation graph, and an audit map to let auditors verify that each step of an LLM‑agent wo…