3 papers
cs.CL2026
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
Elias Lumer, Faheem Nizar, Akshaya Jangiti +4
Recent advancements in Large Language Model (LLM) agents have enabled complex multi-turn agentic tasks requiring extensive tool calling, where conversations can span dozens of API…
cs.CL2025
Jackal: A Real-World Execution-Based Benchmark Evaluating Large Language Models on Text-to-JQL Tasks
Kevin Frank, Anmol Gulati, Elias Lumer +2
Enterprise teams rely on the Jira Query Language (JQL) to retrieve and filter issues from Jira. Yet, to our knowledge, there is no open, real-world, execution-based benchmark for m…
cs.CL2025
Simulation Agent: A Framework for Integrating Simulation and Large Language Models for Enhanced Decision-Making
Jacob Kleiman, Kevin Frank, Joseph Voyles +1
Simulations, although powerful in accurately replicating real-world systems, often remain inaccessible to non-technical users due to their complexity. Conversely, large language mo…