6 papers
WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata
Basel Shbita, Pengyuan Li, Anna Lisa Gentile
Visual Question Answering (VQA) benchmarks have largely emphasized perception-based tasks that can be solved from visual content alone. In contrast, many real-world scenarios requi…
A Systematic Approach for Large Language Models Debugging
Basel Shbita, Anna Lisa Gentile, Bing Zhang +10
Large language models (LLMs) have become central to modern AI workflows, powering applications from open-ended text generation to complex agent-based reasoning. However, debugging…
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
Kellen Tan Cheng, Anna Lisa Gentile, Chad DeLuca +1
The pervasiveness of large language models (LLMs) in enterprise settings has also brought forth a significant amount of risks associated with their usage. Guardrails technologies a…
OneShield -- the Next Generation of LLM Guardrails
Chad DeLuca, Anna Lisa Gentile, Shubhi Asthana +7
The rise of Large Language Models has created a general excitement about the great potential for a myriad of applications. While LLMs offer many possibilities, questions about safe…
Towards Computer-Using Personal Agents
Piero A. Bonatti, John Domingue, Anna Lisa Gentile +9
Computer-Using Agents (CUA) enable users to automate increasingly-complex tasks using graphical interfaces such as browsers. As many potential tasks require personal data, we propo…
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
Shubhi Asthana, Bing Zhang, Ruchi Mahindru +3
The adoption of Large Language Models (LLMs) has revolutionized AI applications but poses significant challenges in safeguarding user privacy. Ensuring compliance with privacy regu…