6 papers
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Logan Ritchie, Sushant Mehta, Liudas Panavas +1
Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…
Cross-Benchmark Generalization in Long-Horizon Agents
Sushant Mehta, Logan Ritchie, Liudas Panavas +1
For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Liudas Panavas, Sebastian Minus, Bradley Monton +4
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Sushant Mehta, Liudas Panavas, Suhaas Garre +1
Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…
Set Visualizations for Comparing and Evaluating Machine Learning Models
Liudas Panavas, Tarik Crnovrsanin, Racquel Fygenson +4
Machine learning practitioners often need to compare multiple models to select the best one for their application. However, current methods of comparing models fall short because t…
But Can You Use It? Design Recommendations for Differentially Private Interactive Systems
Liudas Panavas, Joshua Snoke, Erika Tyagi +2
Accessing data collected by federal statistical agencies is essential for public policy research and improving evidence-based decision making, such as evaluating the effectiveness…