6 papers
Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations
Michael Levit, Josh Ledgard, Haoyu Dong +5
LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representativ…
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar +6
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agent…
Office Comprehension Benchmark
Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…
TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
Reza Esfandiarpoor, Vishwas Suryanarayanan, Stephen H. Bach +2
Since the introduction of the Model Context Protocol (MCP), the number of available tools for Large Language Models (LLMs) has increased significantly. These task-specific tool set…
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar +3
Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious…
Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment
Minh-Quan Le, Gaurav Mittal, Tianjian Meng +5
While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Que…