most citedPhoton: Federated LLM Pre-Training

1 citations · 1 across the 6 of their papers we have counts for

collaborators

27 papers

cs.AI2026

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

James Elcock, William F. Shen, Xinchi Qiu +1

Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment,…

cs.LG2026

The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

Alex Iacob, Andrej Jovanović, William F. Shen +10

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a…

cs.LG2026

FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs

Lorenzo Sani, Zeyu Cao, Meghdad Kurmanji +5

Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. Mixture-of-Experts (MoEs) architectures partially…

cs.LG2026

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

Andrej Jovanović, Alex Iacob, Mher Safaryan +6

Distributed training of foundation models via is limited by interconnect bandwidth. While infrequent communication strategies reduce synchronization frequency, they…

cs.DC2026

PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity

Muhammad Waseem, Nurbek Tastan, Andrej Jovanovic +4

Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware res…

cs.LG20261 cited

Photon: Federated LLM Pre-Training

Lorenzo Sani, Alex Iacob, Zeyu Cao +8

Scaling large language models (LLMs) demands extensive data and computing resources, which are traditionally constrained to data centers by the high-bandwidth requirements of distr…