#agentic systems
11 papers · 1 filter
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
Phuc Pham, Truong-Son Hy
The paper proposes a Function–Evidence–Validation (FEV) framework to evaluate bioinformatics workflows generated by large language model agents, emphasizing workflow correctness an…
Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
Xu Zheng, Zhuomin Chen, Chaohao Lin +4
The paper introduces Trajectory Graph Copilot, a framework that builds probabilistic graphs of past agent trajectories and uses a graph neural network to flag potentially erroneous…
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
Musa Shams
The paper presents LayerRAG-Bench, a benchmark that evaluates the reliability of agentic retrieval-augmented generation systems across multiple layers such as evidence, tool contra…
Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
Soham Gadgil, David Alexander, Sai Sunku +1
The paper investigates how malicious instructions embedded in persistent memory files can be used to launch prompt injection attacks on agentic AI systems, evaluating several large…
CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models
Zhu Cheng, Zhenming Wang, Yu +14
The paper presents CatalogAgent, an agentic system that uses a Supervisor Agent to resolve conflicts between LLM-based generators and evaluators for filling missing product attribu…
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare
The paper presents an explainable, agentic system that uses summary‑based memory to detect multi‑turn conversational scams, introduces a new benchmark dataset (ConScamBench‑278), a…