Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
Spandan Garg, Vikram Nitin, Yufan Huang
Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow responsibilities like search, debugging or…
cs.AI2025
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
Shanchao Liang, Spandan Garg, Roshanak Zilouchian Moghaddam
As large language models (LLMs) become increasingly capable and widely adopted, benchmarks play a central role in assessing their practical utility. For example, SWE-Bench Verified…
cs.AI2025
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
Dhruv Gautam, Spandan Garg, Jinu Jang +2
Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better unde…