5 papers · 1 filter
CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search
Sriram Selvam, Anneswa Ghosh
When several retrieved sources support the same claim, an answer engine cites some but not others. We call this decision citation allocation and introduce CITECHOICE, a causal audi…
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG
Sriram Selvam, Anneswa Ghosh
LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to id…
Similar Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents
Sriram Selvam, Anneswa Ghosh
Search APIs expose ranked snippets, URLs, and metadata on which agents decide whether to answer, search again, or fetch pages. We evaluate these interfaces as decision surfaces on…
ProfileFoundry: A Synthetic Person-Object Substrate for Privacy, Memory, and Tool-Use Evaluation in LLM Agent
Sriram Selvam, Anneswa Ghosh
Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user d…
PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs
Sriram Selvam, Anneswa Ghosh
The memorization of sensitive and personally identifiable information (PII) by large language models (LLMs) poses growing privacy risks as models scale and are increasingly deploye…