benchmarking 1code editing automation 1execution efficiency 1LLM agents 1task complexity estimation 1
From the 1 of 6 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Tmax: A simple recipe for terminal agents
Hamish Ivison, Junjie Oscar Yin, Rulin Shao +3
Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite their prevalence, relatively little academic work has examined R…
cs.CL2025
Approximating Language Model Training Data from Weights
John X. Morris, Junjie Oscar Yin, Woojeong Kim +2
Modern language models often have open weights but closed training data. We formalize the problem of data approximation from model weights and propose several baselines and metrics…