Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
Nicholas Sadjoli, Tim Siefken, Atin Ghosh +2
Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice…
cs.AI2026
Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
Penny Chong, Harshavardhan Abichandani, Jiyuan Shen +4
Agent applications are increasingly adopted to automate workflows across diverse tasks. However, due to the heterogeneous domains they operate in, it is challenging to create a sca…