3 papers
cs.SE2026
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
Bassam Adnan, Aviral Gupta, Sreemaee Akshathala +1
Benchmarks for large language models (LLMs) have progressed from snippet-level function generation to repository-level issue resolution, yet they overwhelmingly target implementati…
cs.MA2025
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
Sreemaee Akshathala, Bassam Adnan, Mahisha Ramesh +3
Recent advances in agentic AI have shifted the focus from standalone Large Language Models (LLMs) to integrated systems that combine LLMs with tools, memory, and other agents to pe…
cs.SE2025
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
Kannan Parthasarathy, Karthik Vaidhyanathan, Rudra Dhar +8
Cloud Operations (CloudOps) is a rapidly growing field focused on the automated management and optimization of cloud infrastructure which is essential for organizations navigating…