activity
20242026
collaborators

6 papers

cs.AI2026

Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer

Logan Ritchie, Sushant Mehta, Liudas Panavas +1

Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…

cs.SE2026

Cross-Benchmark Generalization in Long-Horizon Agents

Sushant Mehta, Logan Ritchie, Liudas Panavas +1

For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…

cs.AI2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Liudas Panavas, Sebastian Minus, Bradley Monton +4

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…

cs.AI2026

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Sushant Mehta, Liudas Panavas, Suhaas Garre +1

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…

cs.HC2025

Set Visualizations for Comparing and Evaluating Machine Learning Models

Liudas Panavas, Tarik Crnovrsanin, Racquel Fygenson +4

Machine learning practitioners often need to compare multiple models to select the best one for their application. However, current methods of comparing models fall short because t…

cs.HC2024

But Can You Use It? Design Recommendations for Differentially Private Interactive Systems

Liudas Panavas, Joshua Snoke, Erika Tyagi +2

Accessing data collected by federal statistical agencies is essential for public policy research and improving evidence-based decision making, such as evaluating the effectiveness…