12 papers
Recursive Harness Self-Improvement
Hyunin Lee, Jinglue Xu, Jeffrey Seely +3
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This…
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Yujin Tang, Chenming Shang, Ruize Xu +1
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to rem…
Learning to Orchestrate Agents in Natural Language with the Conductor
Stefan Nielsen, Edoardo Cetin, Peter Schwendeman +3
Powerful large language models (LLMs) from different providers have been expensively trained and finetuned to specialize across varying domains. In this work, we introduce a new ki…
TRINITY: An Evolved LLM Coordinator
Jinglue Xu, Qi Sun, Peter Schwendeman +3
Combining diverse foundation models is promising, but weight-merging is limited by mismatched architectures and closed APIs. Trinity addresses this with a lightweight coordinator t…
Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
Yifan Hu, Jie Yang, Tian Zhou +4
Although contrastive and other representation-learning methods have long been explored in vision and NLP, their adoption in modern time series forecasters remains limited. We belie…
Evolutionary Context Search for Automated Skill Acquisition
Qi Sun, Stefan Nielsen, Rio Yokota +1
Large Language Models cannot reliably acquire new knowledge post-deployment -- even when relevant text resources exist, models fail to transform them into actionable knowledge with…