From the 9 of 111 papers with an AI index.
38 citations
- Australian Regenerative Medicine InstituteAU39 papers
- Massachusetts Institute of TechnologyUS37 papers
- University of ZurichCH37 papers
- Eötvös Loránd UniversityHU36 papers
- Rutherford Appleton LaboratoryGB36 papers
- Université Paris-SaclayFR36 papers
- University of Chinese Academy of SciencesCN36 papers
- University of Maryland, College ParkUS36 papers
- Istituto Nazionale di Fisica Nucleare, Sezione di BolognaIT35 papers
- Jagiellonian UniversityPL35 papers
- Peking UniversityCN35 papers
- Tsinghua UniversityCN35 papers
4 papers · 1 filter
AgentChaos: Chaos Engineering for Agent Systems via Programmatic Fault Injection
Gou Tan, Zhensu Sun, Jieke Shi +10
Agent systems rely on LLM APIs for every response, but these APIs can return server errors, truncated responses, or corrupted content that propagates through downstream agents and…
Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering
Yunpeng Xiong, Ting Zhang
In this paper, we present a comparative study of three state-of-the-art LLM-based agent frameworks, i.e., Aider, OpenHands, and SWE-agent, for vulnerability FP filtering. We evalua…
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
Aaron Guoxiang Guo, Aldeida Aleti, Neelofar Neelofar +3
With the widespread application of LLM-based dialogue systems in daily life, quality assurance has become more important than ever. Recent research has successfully introduced meth…
HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
Kla Tantithamthavorn, Hong Yi Lin, Patanamon Thongtanunam +3
Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated…