6 papers
Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations
Michael Levit, Josh Ledgard, Haoyu Dong +5
LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representativ…
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar +6
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agent…
Office Comprehension Benchmark
Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…
TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
Reza Esfandiarpoor, Vishwas Suryanarayanan, Stephen H. Bach +2
Since the introduction of the Model Context Protocol (MCP), the number of available tools for Large Language Models (LLMs) has increased significantly. These task-specific tool set…
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar +3
Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious…
Local Prompt Optimization
Yash Jain, Vishal Chowdhary
In recent years, the use of prompts to guide the output of Large Language Models have increased dramatically. However, even the best of experts struggle to choose the correct words…