collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau +1

Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on…

cs.AI2026

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…

cs.AI2026

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel +1

Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks…

cs.AI2026

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Alimurtaza Mustafa Merchant, Krish Veera, Sajal Kumar Goyla +3

Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure modes, forecasting tools, an…

cs.AI2026

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

Yusheng Li, Tianjun Feng, Yunfeng Chen +6

LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critic…