1 citations · 1 across the 3 of their papers we have counts for
4 papers
Tangent: An Empirical Study of Testing Practices for LLM-Based Agent Applications
Rangeet Pan, Tyler Stennett, Divya Sankar +5
Agents built on large language models (LLMs) are increasingly used to build applications that perform complex, multi-step tasks involving reasoning, tool use, and interaction with…
Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions
Tyler Stennett, Rangeet Pan, Bridget McGinn +2
Research on automating software testing has spanned several decades. Most existing approaches generate unit tests for individual methods, validate isolated API endpoints, or target…
ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java
Advait Pavuluri, Bridget McGinn, Ashita Saxena +6
Java remains central to enterprise software, and many applications outlive their original architecture. Migrating them across frameworks is a behavior-preserving refactoring spanni…
Scaling Granite Code Models to 128K Context
Matt Stallone, Vaibhav Saxena, Leonid Karlinsky +19
This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code mo…