activity
20242026
collaborators
Showing cs.AIShow all

12 papers · 1 filter

cs.AI2026

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…

cs.AI2026

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

Kushal Raj Bhandari, Ling Yue, Ching-Yun Ko +4

Compact language models (LMs) reduce cost, latency, and deployment risk for tool agents. Yet MCP-style tool use requires more than isolated function calling: an agent must discover…

cs.AI2026

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel +1

Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks…

cs.AI2026

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Harshada Badave, Santosh Borse, Andrea Gomez +6

Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate on…

cs.AI2026

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Alimurtaza Mustafa Merchant, Krish Veera, Sajal Kumar Goyla +3

Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure modes, forecasting tools, an…

cs.AI2026

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

Yusuke Ozaki, Dhaval Patel

Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long workflows, leading to brittle fa…