14 papers
Sakana Fugu Technical Report
Yujin Tang, Edoardo Cetin, Jinglue Xu +11
The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next ob…
Sparser, Faster, Lighter Transformer Language Models
Edoardo Cetin, Stefano Peluchetti, Emilio Castillo +3
Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, we tackle these costs by leveraging uns…
Learning to Orchestrate Agents in Natural Language with the Conductor
Stefan Nielsen, Edoardo Cetin, Peter Schwendeman +3
Powerful large language models (LLMs) from different providers have been expensively trained and finetuned to specialize across varying domains. In this work, we introduce a new ki…
TRINITY: An Evolved LLM Coordinator
Jinglue Xu, Qi Sun, Peter Schwendeman +3
Combining diverse foundation models is promising, but weight-merging is limited by mismatched architectures and closed APIs. Trinity addresses this with a lightweight coordinator t…
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
Gianluigi Silvestri, Edoardo Cetin
Reasoning-oriented language models achieve strong performance by generating long chain-of-thought traces at inference time. However, this capability comes with substantial and ofte…
Doc-to-LoRA: Learning to Instantly Internalize Contexts
Rujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka +1
Long input sequences are central to in-context learning, document understanding, and multi-step reasoning of Large Language Models (LLMs). However, the quadratic attention cost of…