3 citations · 3 across the 1 of their papers we have counts for
3 papers
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
Prompt Perturbation Consistency Learning for Robust Language Models
Yao Qiang, Subhrangshu Nandi, Ninareh Mehrabi +4
Large language models (LLMs) have demonstrated impressive performance on a number of natural language processing tasks, such as question answering and text summarization. However,…
Measuring and Mitigating Local Instability in Deep Neural Networks
Arghya Datta, Subhrangshu Nandi, Jingcheng Xu +4
Deep Neural Networks (DNNs) are becoming integral components of real world services relied upon by millions of users. Unfortunately, architects of these systems can find it difficu…