collaborators

5 papers

cs.CL2026

ClaimDB: A Fact Verification Benchmark over Large Structured Data

Michael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah +1

Real-world fact-checking often involves verifying claims grounded in structured data at scale. Despite substantial progress in fact-verification benchmarks, this setting remains la…

cs.CL2026

iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics

Preetam Prabhu Srikar Dammu, Arnav Palkhiwala, Tanya Roosta +1

With the emergence of search-enabled generative QA systems, users are increasingly turning to tools that browse, aggregate, and reconcile evidence across multiple sources on their…

cs.IR2025

Beyond Static Evaluation: Rethinking the Assessment of Personalized Agent Adaptability in Information Retrieval

Kirandeep Kaur, Preetam Prabhu Srikar Dammu, Hideo Joho +1

Personalized AI agents are becoming central to modern information retrieval, yet most evaluation methodologies remain static, relying on fixed benchmarks and one-off metrics that f…

cs.IR2025

Dynamic Evaluation Framework for Personalized and Trustworthy Agents: A Multi-Session Approach to Preference Adaptability

Chirag Shah, Hideo Joho, Kirandeep Kaur +1

Recent advancements in generative AI have significantly increased interest in personalized agents. With increased personalization, there is also a greater need for being able to tr…

cs.CL2025

Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering Datasets

Preetam Prabhu Srikar Dammu, Himanshu Naidu, Chirag Shah

As question answering (QA) systems advance alongside the rapid evolution of foundation models, the need for robust, adaptable, and large-scale evaluation benchmarks becomes increas…