4 papers
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
Alimurtaza Mustafa Merchant, Krish Veera, Sajal Kumar Goyla +3
Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure modes, forecasting tools, an…
MuseScorer: Idea Originality Scoring At Scale
Ali Sarosh Bangash, Krish Veera, Ishfat Abrar Islam +1
An objective, face-valid method for scoring idea originality is to measure each idea's statistical infrequency within a population -- an approach long used in creativity research.…
AI Can Enhance Creativity in Social Networks
Raiyan Abdul Baten, Ali Sarosh Bangash, Krish Veera +2
Can peer recommendation engines elevate people's creative performances in self-organizing social networks? Answering this question requires resolving challenges in data collection…