4 papers
Office Comprehension Benchmark
Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar +6
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agent…
A MISMATCHED Benchmark for Scientific Natural Language Inference
Firoz Shaik, Mobashir Sadat, Nikita Gautam +2
Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this…
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar +3
Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious…