works on

From the 1 of 7 linked papers with an AI index.

most citedEnhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms

2 citations · 2 across the 4 of their papers we have counts for

collaborators

7 papers

cs.LG2026

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Hei Yi Mak, Shadan Golestan, Hoang Le +10

The paper introduces HiFloat4, a 4-bit floating-point format and a Rollout Residual Quantization technique that enable end-to-end reinforcement learning post‑training of large lang…

cs.LG20262 cited

Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms

Sheila Schoepp, Mehran Taghian, Shotaro Miwa +3

Industry is moving toward autonomous, network-connected machines that detect and adapt to changing conditions, including hardware faults. Conventional fault-tolerant design duplica…

cs.LG2026

HiFloat4 Format for Language Model Pre-training on Ascend NPUs

Mehran Taghian, Yunke Peng, Xing Huang +22

Large foundation models have become central to modern machine learning, with performance scaling predictably with model size and data. However, training and deploying such models i…

cs.LG2026

WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation

Zahin Sufiyan, Shadan Golestan, Yoshihiro Mitsuka +2

Generative Flow Networks for continuous scenarios (CFlowNets) have shown promise in solving sequential decision-making tasks by learning stochastic policies using a flow and a retr…

cs.LG2026

TLXML: Task-Level Explanation of Meta-Learning via Influence Functions

Yoshihiro Mitsuka, Shadan Golestan, Zahin Sufiyan +2

Meta-learning enables models to rapidly adapt to new tasks by leveraging prior experience, but its adaptation mechanisms remain opaque, especially regarding how past training tasks…

cs.LG2025

The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning

Sheila Schoepp, Masoud Jafaripour, Yingyue Cao +6

Reinforcement learning (RL) has shown impressive results in sequential decision-making tasks. Meanwhile, Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged…