4 papers
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan +6
Reinforcement Learning with Verified Reward (RLVR) has emerged as a critical paradigm for advancing the reasoning capabilities of Large Language Models (LLMs). Most existing RLVR m…
QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL
Cong Yu, Valter Uotila, Shilong Deng +5
Designing and optimizing task-specific quantum circuits are crucial to leverage the advantage of quantum computing. Recent large language model (LLM)-based quantum circuit generati…
Improving the Transferability of Adversarial Attacks by an Input Transpose
Qing Wan, Shilong Deng, Xun Wang
Deep neural networks (DNNs) are highly susceptible to adversarial examples--subtle perturbations applied to inputs that are often imperceptible to humans yet lead to incorrect mode…
From Natural Language to Extensive-Form Game Representations
Shilong Deng, Yongzhao Wang, Rahul Savani
We introduce a framework for translating game descriptions in natural language into extensive-form representations in game theory, leveraging Large Language Models (LLMs) and in-co…