activity
20242026
collaborators

6 papers

cs.AI2026

NAAMSE: Framework for Evolutionary Security Evaluation of Agents

Kunal Pai, Parth Shah, Harshil Patel

AI agents are increasingly deployed in production, yet their security evaluations remain bottlenecked by manual red-teaming or static benchmarks that fail to model adaptive, multi-…

cs.CL2025

Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III

Pranam Shetty, Abhisek Upadhayaya, Parth Mitesh Shah +3

As financial institutions increasingly adopt Large Language Models (LLMs), rigorous domain-specific evaluation becomes critical for responsible deployment. This paper presents a co…

cs.AI2025

How Many Instructions Can LLMs Follow at Once?

Daniel Jaroslawicz, Brendan Whiting, Parth Shah +1

Production-grade LLM systems require robust adherence to dozens or even hundreds of instructions simultaneously. However, the instruction-following capabilities of LLMs at high ins…

cs.MA2025

HASHIRU: Hierarchical Agent System for Hybrid Intelligent Resource Utilization

Kunal Pai, Parth Shah, Harshil Patel

Rapid Large Language Model (LLM) advancements are fueling autonomous Multi-Agent System (MAS) development. However, current frameworks often lack flexibility, resource awareness, m…

cs.CL2025

GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs

Longchao Da, Parth Mitesh Shah, Kuan-Ru Liou +2

Large Language Models are now key assistants in human decision-making processes. However, a common note always seems to follow: "LLMs can make mistakes. Be careful with important i…

cs.CL2024

On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models

April Yang, Jordan Tab, Parth Shah +1

The increasing reliance on large language models (LLMs) for diverse applications necessitates a thorough understanding of their robustness to adversarial perturbations and out-of-d…