3 papers
cs.CL2026
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie +2
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop…
cs.IR2026
When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
Syed Shariyar Murtaza, Yifan Nie, Utkarsh Soni +2
LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over…
cs.AI2026
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti +4
Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-…