Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
Jacob Morrison, Sanjay Adhikesaven, Akshita Bhagia +3
Extending a fully post-trained language model with new domain capabilities is fundamentally limited by monolithic training paradigms: retraining from scratch is expensive and scale…
cs.LG2026
Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization
Joesph An, Joseph An, Phillip Keung +3
Audio language models process input audio into rich frame-level representations, but the standard approach to temporal localization generates timestamps as sequences of text tokens…