Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Diffract: Spectral View of LLM Domain Adaptation
Nikita Borodin, Maria Krylova, Artem Zabolotnyi +6
We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Us…
cs.LG2026
Never Skip a Batch: Dense Learning of Temporal GNNs via Adaptive Pseudo-Supervision
Alexander Panyshev, Dmitry Vinichenko, Oleg Travkin +2
Temporal graph networks suffer from irregular supervision in realworld dynamic graphs, as most minibatches contain few labeled events. The lack of labels leads to high-variance gra…