#llm agents
49 papers · 1 filter
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Albert Gong, Kyuseong Choi, Abhineet Agarwal +5
The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…
MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
Dongyi Liu, Haixing He, Xiaobao Wu +1
The paper introduces MIND, a lightweight framework that uses an intent‑aware information bottleneck to detect and filter poisoned memory in large language model agents, reducing at…
ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents
Stef Cuyckens, Mihaela Jivanescu, Jun Yin +2
The paper presents ARES, a framework that adaptively controls the reasoning effort of large language model agents when optimizing RTL designs for power, performance, and area, whil…
ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
Yongye Su, Wujiang Xu, Chaoji Zuo +1
ChronoMem adds a semantic version‑control layer to large language model agents, allowing them to snapshot, browse, and roll back their long‑term memory using natural‑language reque…
Baikal: Structured Search for Deep Research over Data Lakes
Dhruv Agarwal, Rishitha Guttapalle Mohan, Aarti Kumari +5
Baikal is a framework that clusters heterogeneous tables and passages into semantic regions and uses adaptive, budgeted search policies to guide an LLM agent in generating subquest…
Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
Zuyuan Zhang, Hanqing Yang, Carlee Joe-Wong +1
The paper proposes iCORE, a unified representation that combines a cooperation graph, an obligation graph, and an audit map to let auditors verify that each step of an LLM‑agent wo…