works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.LG2026

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

Yassine Chemingui, Chenhua Fan, Honghao Wei +1

SteinGate introduces a distributional safety certificate based on Kernelized Stein Discrepancy to detect rare, high-cost tail events in reinforcement learning and dynamically switc…

cs.LG2026

Conformal Margin Risk Minimization: An Envelope Framework for Robust Learning under Label Noise

Yuanjie Shi, Peihong Li, Zijian Zhang +2

Most methods for learning with noisy labels require privileged knowledge such as noise transition matrices, clean subsets or pretrained feature extractors, resources typically unav…

cs.AI2026

Discovery of Feasible 3D Printing Configurations for Metal Alloys via AI-driven Adaptive Experimental Design

Azza Fadhel, Nathaniel W. Zuckschwerdt, Aryan Deshwal +3

Configuring the parameters of additive manufacturing processes for metal alloys is a challenging problem due to complex relationships between input parameters (e.g., laser power, s…

cs.SE2025

An Exploratory Study of Bayesian Prompt Optimization for Test-Driven Code Generation with Large Language Models

Shlok Tomar, Aryan Deshwal, Ethan Villalovoz +3

We consider the task of generating functionally correct code using large language models (LLMs). The correctness of generated code is influenced by the prompt used to query the giv…

cs.LG2025

Online Optimization for Offline Safe Reinforcement Learning

Yassine Chemingui, Aryan Deshwal, Alan Fern +2

We study the problem of Offline Safe Reinforcement Learning (OSRL), where the goal is to learn a reward-maximizing policy from fixed data under a cumulative cost constraint. We pro…

cs.LG2025

Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning

Yassine Chemingui, Aryan Deshwal, Honghao Wei +2

Offline safe reinforcement learning (OSRL) involves learning a decision-making policy to maximize rewards from a fixed batch of training data to satisfy pre-defined safety constrai…