4 papers
Rate or Fate? RLVR: Reinforcement Learning with Verifiable Noisy Rewards
Ali Rad, Khashayar Filom, Darioush Keivan +2
Reinforcement learning with verifiable rewards (RLVR) is a simple but powerful paradigm for training LLMs: sample a completion, verify it, and update. In practice, however, the ver…
TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
Bardia Nadimi, Khashayar Filom, Deming Chen +1
With the rapid advancement of Large Language Models (LLMs), there is growing interest in applying them to hardware design and verification. Among these stages, design verification…
MBExplainer: Multilevel bandit-based explanations for downstream models with augmented graph embeddings
Ashkan Golgoon, Ryan Franks, Khashayar Filom +1
In many industrial applications, it is common that the graph embeddings generated from training GNNs are used in an ensemble model where the embeddings are combined with other tabu…
Mechanistic interpretability of large language models with applications to the financial services industry
Ashkan Golgoon, Khashayar Filom, Arjun Ravi Kannan
Large Language Models such as GPTs (Generative Pre-trained Transformers) exhibit remarkable capabilities across a broad spectrum of applications. Nevertheless, due to their intrins…