6 papers
Recursive Harness Self-Improvement
Hyunin Lee, Jinglue Xu, Jeffrey Seely +3
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This…
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning
Jiayi Lei, Yuandong Pu, Xingyu Han +8
Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their s…
How Good Can Linear Models Be for Time-Series Forecasting?
Lang Huang, Jinglue Xu, Luke Darlow
Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers to general-purpose foundation models, on the assumption that ca…
Sakana Fugu Technical Report
Yujin Tang, Edoardo Cetin, Jinglue Xu +11
The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next ob…
Learning to Orchestrate Agents in Natural Language with the Conductor
Stefan Nielsen, Edoardo Cetin, Peter Schwendeman +3
Powerful large language models (LLMs) from different providers have been expensively trained and finetuned to specialize across varying domains. In this work, we introduce a new ki…
TRINITY: An Evolved LLM Coordinator
Jinglue Xu, Qi Sun, Peter Schwendeman +3
Combining diverse foundation models is promising, but weight-merging is limited by mismatched architectures and closed APIs. Trinity addresses this with a lightweight coordinator t…