works on

From the 2 of 16 linked papers with an AI index.

collaborators

13 papers

cs.LG2026

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Hei Yi Mak, Shadan Golestan, Hoang Le +10

The paper introduces HiFloat4, a 4-bit floating-point format and a Rollout Residual Quantization technique that enable end-to-end reinforcement learning post‑training of large lang…

cs.LG2026

Stable FP4 Training via Transposition-Invariant Block Quantization

Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi +6

Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging…

cs.AI2026

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

Sijie Wang, Zhengyu Qing, Zhiqiang Tan +6

Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models…

astro-ph.GA2026

HOTDISK. Finding Massive Protostellar Disks with Water and Refractory Molecular Species

Kai Yang, Yichen Zhang, Kei E. I. Tanaka +16

We present high-angular-resolution (, au) ALMA Band~6 observations from the HOTDISK project (Hot-Origin Tracer survey of DISKs of massive pro…

cs.LG2026

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training

Boao Kong, Weichen Jia, Engao Zhang +6

Training stability is a key bottleneck in low-precision language model training: efficient low-cost paths can still produce short-lived numerical risks at a small set of operators.…

cs.DC2026

TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation

Huichao Chai, Zhixin Wu, Xuemiao Li +8

Generative recommendation (GR) has emerged as a promising paradigm that replaces fragmented, scenario-specific architectures with unified Transformer-based models, exhibiting scali…