collaborators

5 papers

cs.LG2026

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar +6

Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agent…

cs.CL2026

Office Comprehension Benchmark

Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17

We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…

cs.CL2025

Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper

Krishna Garg, Firoz Shaik, Sambaran Bandyopadhyay +1

As researchers increasingly adopt LLMs as writing assistants, generating high-quality research paper introductions remains both challenging and essential. We introduce Scientific I…

cs.LG2025

VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents

Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar +3

Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious…

cs.CL2025

A MISMATCHED Benchmark for Scientific Natural Language Inference

Firoz Shaik, Mobashir Sadat, Nikita Gautam +2

Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this…