3 papers
cs.AI2026
Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse
Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan
Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task complet…
cs.CL2026
Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation
Devansh Singh, Sundaraparipurnan Narayanan
Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A…
cs.CL2025
Unmasking the Reality of PII Masking Models: Performance Gaps and the Call for Accountability
Devansh Singh, Sundaraparipurnan Narayanan
Privacy Masking is a critical concept under data privacy involving anonymization and de-anonymization of personally identifiable information (PII). Privacy masking techniques rely…