3 papers
cs.CL2026
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
Aashna Garg, Siddharth Singha Roy, Jinu Jang +2
Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary strong-vs-weak decisions and c…
cs.SE2025
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Avi Arora, Jinu Jang, Roshanak Zilouchian Moghaddam
Modern Large Language Model (LLM) agents promise end to end assistance with real-world software tasks, yet existing benchmarks evaluate LLM agents almost exclusively in pre-baked e…
cs.AI2025
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
Dhruv Gautam, Spandan Garg, Jinu Jang +2
Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better unde…