4 papers
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Youssef Mroueh, Nicolas Dupuis, Brian Belgodere +6
We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Opti…
Customizing a Large Language Model for VHDL Design of High-Performance Microprocessors
Nicolas Dupuis, Ravi Nair, Shyam Ramji +7
The use of Large Language Models (LLMs) in hardware design has taken off in recent years, principally through its incorporation in tools that increase chip designer productivity. T…
Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models
Sanjay Vishwakarma, Francis Harkins, Siddharth Golecha +7
Quantum programs are typically developed using quantum Software Development Kits (SDKs). The rapid advancement of quantum computing necessitates new tools to streamline this develo…
Automated Code generation for Information Technology Tasks in YAML through Large Language Models
Saurabh Pujar, Luca Buratti, Xiaojie Guo +8
The recent improvement in code generation capabilities due to the use of large language models has mainly benefited general purpose programming languages. Domain specific languages…