5 papers
Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation
Hong Chen, Pengcheng Wu, Yuanguo Lin +4
We rethink Federated Learning (FL) from a nested learning perspective, framing the core challenge as how to collaboratively learn optimization rules, not just static models, to tac…
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
Runzhe Zhang, Letian Chen, Wenpeng Zhang +2
We present FlowLM, a flow matching language model transformed from pre-trained diffusion language models via efficient fine-tuning. By re-aligning the curved sampling trajectories…
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
Yangyi Fang, Jiaye Lin, Xiaoliang Fu +4
Multi-turn LLM agents are becoming pivotal to production systems, spanning customer service automation, e-commerce assistance, and interactive task management, where accurately dis…
SeqPE: Transformer with Sequential Position Encoding
Huayang Li, Yahui Liu, Hongyu Sun +5
Since self-attention layers in Transformers are permutation invariant by design, positional encodings must be explicitly incorporated to enable spatial understanding. However, fixe…
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
Shansan Gong, Shivam Agarwal, Yizhe Zhang +9
Diffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, c…