4 papers
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
Kristen Pereira, Neelabh Sinha, Rajat Ghosh +1
Recent advances in frontier large language models have enabled code review agents that operate in open-ended, reasoning-intensive settings. However, the lack of standardized benchm…
QA-prompting: Improving Summarization with Large Language Models using Question-Answering
Neelabh Sinha
Language Models (LMs) have revolutionized natural language processing, enabling high-quality text generation through prompting and in-context learning. However, models often strugg…
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
Neelabh Sinha, Vinija Jain, Aman Chadha
The rapid rise of Language Models (LMs) has expanded their use in several applications. Yet, due to constraints of model size, associated cost, or proprietary restrictions, utilizi…
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
Neelabh Sinha, Vinija Jain, Aman Chadha
Visual Question-Answering (VQA) has become key to user experience, particularly after improved generalization capabilities of Vision-Language Models (VLMs). But evaluating VLMs for…