activity
20242026
collaborators

5 papers

cs.AI2026

VeRO: A Harness for Agents to Optimize Agents

Varun Ursekar, Apaar Shanker, Veronica Chatrath +2

An important emerging application of coding agents is agent harness optimization: the iterative improvement of a target agent by editing and evaluating its code. Despite its releva…

cs.IR2026

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

Andrew Klearman, Radu Revutchi, Rohin Garg +3

Retrieval quality is the primary bottleneck for accuracy and robustness in retrieval-augmented generation (RAG). Current evaluation relies on heuristically constructed query sets,…

cs.CL2026

LHAW: Controllable Underspecification for Long-Horizon Tasks

George Pu, Michael S. Lee, Udari Madhushani Sehwag +6

Long-horizon workflow agents that operate effectively over extended periods are essential for truly autonomous systems. Their reliable execution critically depends on the ability t…

cs.CL2025

Assessing Robustness to Spurious Correlations in Post-Training Language Models

Julia Shuieh, Prasann Singhal, Apaar Shanker +3

Supervised and preference-based fine-tuning techniques have become popular for aligning large language models (LLMs) with user intent and correctness criteria. However, real-world…

cs.CL2024

Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs

Yung-Chieh Chan, George Pu, Apaar Shanker +4

As large language models (LLMs) are applied to more use cases, creating high quality, task-specific datasets for fine-tuning becomes a bottleneck for model improvement. Using high…