1 citations · 1 across the 1 of their papers we have counts for
4 papers
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
Shahin Honarvar, Amber Gorzynski, James Lee-Jones +4
Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing pointwise benchmarks offer limited insi…
TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
Jacob Si, Mike Qu, Michelle Lee +2
Incorporating external knowledge bases in traditional retrieval-augmented generation (RAG) relies on parsing the document, followed by querying a language model with the parsed inf…
Fine-tuning with RAG for Improving LLM Learning of New Skills
Humaid Ibrahim, Nikolai Rozanov, Marek Rei
Large language model (LLM) agents deployed for multi-step tasks frequently fail in predictable ways: attempting actions with unmet preconditions, issuing redundant commands, or mis…
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
Nikolai Rozanov, Marek Rei
Large language models (LLMs) are increasingly used as autonomous agents, tackling tasks from robotics to web navigation. Their performance depends on the underlying base agent. Exi…