activity
20242026
collaborators

5 papers

cs.SE2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…

cs.SE2025

Robust Learning of Diverse Code Edits

Tushar Aggarwal, Swayam Singh, Abhijeet Awasthi +2

Software engineering activities frequently involve edits to existing code. However, contemporary code language models (LMs) lack the ability to handle diverse types of code-edit re…

cs.CL2025

Language Models' Factuality Depends on the Language of Inquiry

Tushar Aggarwal, Kumar Tanmay, Ayush Agrawal +3

Multilingual language models (LMs) are expected to recall factual knowledge consistently across languages, yet they often fail to transfer knowledge between languages even when the…

cs.CL2025

PASS: Presentation Automation for Slide Generation and Speech

Tushar Aggarwal, Aarohi Bhand

In today's fast-paced world, effective presentations have become an essential tool for communication in both online and offline meetings. The crafting of a compelling presentation…

cs.SE2024

NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness

Manav Singhal, Tushar Aggarwal, Abhijeet Awasthi +2

Existing evaluation benchmarks of language models of code (code LMs) focus almost exclusively on whether the LMs can generate functionally-correct code. In real-world software engi…