3 papers
cs.CR2025
ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits
Chen Chen, Zaiyan Xu, Mohamadreza Rostami +4
Processor designs rely on iterative modifications and reuse well-established designs. However, this reuse of prior designs also leads to similar vulnerabilities across multiple pro…
cs.LG2025
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
Zaiyan Xu, Sushil Vemuri, Kishan Panaganti +3
A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, a…
cs.LG2023
Bridging Distributionally Robust Learning and Offline RL: An Approach to Mitigate Distribution Shift and Partial Data Coverage
Kishan Panaganti, Zaiyan Xu, Dileep Kalathil +1
The goal of an offline reinforcement learning (RL) algorithm is to learn optimal polices using historical (offline) data, without access to the environment for online exploration.…