2 citations · 4 across the 6 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
On the "Causality" Step in Policy Gradient Derivations: A Pedagogical Reconciliation of Full Return and Reward-to-Go
Nima H. Siboni
In introductory presentations of policy gradients, one often derives the REINFORCE estimator using the full trajectory return and then states, by ``causality,'' that the full retur…
cs.AI2026
LLM-Driven Heuristic Synthesis for Industrial Process Control: Lessons from Hot Steel Rolling
Nima H. Siboni, Seyedreza Kiamousavi, Emad Scharifi
Industrial process control demands policies that are interpretable and auditable, requirements that black-box neural policies struggle to meet. We study an LLM-driven heuristic syn…