2 papers
cs.CL2026
SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs
Jiacheng Sang, Mengyuan Li, Sanxing Chen +3
LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to ad…
cs.CR2026
AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories
Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis +6
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pr…