papers

Publications (40)

cs.CL2026

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL

Jiakang Wang, Runze Liu, Qingpeng Cai +7

Reinforcement learning (RL) has shown great promise in large language models (LLMs) post-training, which typically rely on token-level clipping to maintain stability during optimiz…

cs.LG2026

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Bingxu Liu, Jiashun Liu, Johan Obando-Ceron +5

While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stat…

cs.LG2026

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

Runze Liu, Jiashun Liu, Xu Wan +2

Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT is expected to provide a usefu…

cs.CL2025

A Survey of Reinforcement Learning for Large Reasoning Models

Kaiyan Zhang, Yuxin Zuo, Bingxiang He +36

In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontie…

cs.CL2025

ReviewRL: Towards Automated Scientific Review with RL

Sihang Zeng, Kai Tian, Kaiyan Zhang +9

Peer review is essential for scientific progress but faces growing challenges due to increasing submission volumes and reviewer fatigue. Existing automated review approaches strugg…

cs.CL2026

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Jiakang Wang, Runze Liu, Fuzheng Zhang +3

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training method for improving the reasoning abilities of Large Language Models (LLMs). However, e…

quant-ph2024

A hybrid single quantum dot coupled cavity on a CMOS-compatible SiC photonic chip for Purcell-enhanced deterministic single-photon emission

Yifan Zhu, Runze Liu, Ailun Yi +10

The ability to control nonclassical light emission from a single quantum emitter by an integrated cavity may unleash new perspectives for integrated photonic quantum applications.…

cs.CV2026

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

Taiting Lu, Kaiyuan Lin, Ziwei Dong +18

Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However, their ability to reason abou…

cs.SD2026

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Ye Lu, Yihan Yan, Zhaoyang Zhang +4

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens s…

eess.IV2025

Individualized Deepfake Detection Exploiting Traces Due to Double Neural-Network Operations

Mushfiqur Rahman, Runze Liu, Chau-Wai Wong +1

In today's digital landscape, journalists urgently require tools to verify the authenticity of facial images and videos depicting specific public figures before incorporating them…

physics.comp-ph2020

RoeNets: Predicting Discontinuity of Hyperbolic Systems from Continuous Data

Shiying Xiong, Xingzhe He, Yunjin Tong +2

We introduce Roe Neural Networks (RoeNets) that can predict the discontinuity of the hyperbolic conservation laws (HCLs) based on short-term discontinuous and even continuous train…

cs.LG2025

PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning

Shengjie Sun, Jiafei Lyu, Runze Liu +4

Offline imitation learning (offline IL) enables training effective policies without requiring explicit reward annotations. Recent approaches attempt to estimate rewards for unlabel…

cs.LG2024

SEABO: A Simple Search-Based Method for Offline Imitation Learning

Jiafei Lyu, Xiaoteng Ma, Le Wan +3

Offline reinforcement learning (RL) has attracted much attention due to its ability in learning from static offline datasets and eliminating the need of interacting with the enviro…

cs.CV2026

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Taiting Lu, Runze Liu, Ziwei Dong +18

Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the…

cs.LG2025

Bohdi: Heterogeneous LLM Fusion with Automatic Data Exploration

Junqi Gao, Zhichang Guo, Dazhi Zhang +5

Heterogeneous Large Language Model (LLM) fusion integrates the strengths of multiple source LLMs with different architectures into a target LLM with low computational overhead. Whi…

cs.LG2020

Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and Memorization

Xiaotao Jia, Jianlei Yang, Runze Liu +3

Bayesian method is capable of capturing real world uncertainties/incompleteness and properly addressing the over-fitting issue faced by deep neural networks. In recent years, Bayes…

cs.LG2024

A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning

Shengjie Sun, Runze Liu, Jiafei Lyu +3

Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often…

eess.SP2019

eSLAM: An Energy-Efficient Accelerator for Real-Time ORB-SLAM on FPGA Platform

Runze Liu, Jianlei Yang, Yiran Chen +1

Simultaneous Localization and Mapping (SLAM) is a critical task for autonomous navigation. However, due to the computational complexity of SLAM algorithms, it is very difficult to…

cs.LG2025

VLP: Vision-Language Preference Learning for Embodied Manipulation

Runze Liu, Chenjia Bai, Jiafei Lyu +3

Reward engineering is one of the key challenges in Reinforcement Learning (RL). Preference-based RL effectively addresses this issue by learning from human feedback. However, it is…

cs.CV2025

Unsupervised Monocular Depth Estimation Based on Hierarchical Feature-Guided Diffusion

Runze Liu, Dongchen Zhu, Guanghui Zhang +5

Unsupervised monocular depth estimation has received widespread attention because of its capability to train without ground truth. In real-world scenarios, the images may be blurry…

cs.CL2025

GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Jian Zhao, Runze Liu, Kaiyan Zhang +8

Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However…

eess.SP2021

On Microstructure Estimation Using Flatbed Scanners for Paper Surface Based Authentication

Runze Liu, Chau-Wai Wong

Paper surfaces under the microscopic view are observed to be formed by intertwisted wood fibers. Such structures of paper surfaces are unique from one location to another and are a…

cs.CV2026

OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout

Taiting Lu, Kaiyuan Lin, Mingjia Wang +12

Recent large language models (LLMs) have demonstrated remarkable progress in 3D spatial reasoning, spatial grounding, and fine-grained geometric understanding. However, their abili…

cond-mat.mtrl-sci2025

A Fast, Accurate, and Reactive Equivariant Foundation Potential

Tsz Wai Ko, Runze Liu, Adesh Rohan Mishra +3

Electrostatics govern charge transfer and reactivity in materials. Yet, most foundation potentials (FPs) either do not explicitly model such interactions or pay a prohibitive scali…

cond-mat.mtrl-sci2025

A Foundational Potential Energy Surface Dataset for Materials

Aaron D. Kaplan, Runze Liu, Ji Qi +6

Accurate potential energy surface (PES) descriptions are essential for atomistic simulations of materials. Universal machine learning interatomic potentials (UMLIPs) offer…

cs.LG2026

Temporal Difference Learning with Constrained Initial Representations

Jiafei Lyu, Jingwen Yang, Zhongjian Qiao +5

Recently, there have been numerous attempts to enhance the sample efficiency of off-policy reinforcement learning (RL) agents when interacting with the environment, including archi…

physics.optics2025

Relayed-loop optical scan amplification

Harishankar Jayakumar, Christopher Warkentin, Deano Farinella +4

Many modern sensing, processing, and fabrication technologies depend upon the sequential scanning of laser light. Due to inertial, thermal, and electrical limitations, the speed of…

cs.LG2021

Modeling the Nonsmoothness of Modern Neural Networks

Runze Liu, Chau-Wai Wong, Huaiyu Dai

Modern neural networks have been successful in many regression-based tasks such as face recognition, facial landmark detection, and image generation. In this work, we investigate a…

cs.ET2019

SPINBIS: Spintronics based Bayesian Inference System with Stochastic Computing

Xiaotao Jia, Jianlei Yang, Pengcheng Dai +3

Bayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-in…

cs.CL2025

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Runze Liu, Junqi Gao, Jian Zhao +5

Test-Time Scaling (TTS) is an important method for improving the performance of Large Language Models (LLMs) by using additional computation during the inference phase. However, cu…

cs.LG2025

Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models

Runze Liu, Jiakang Wang, Yuling Shi +11

Reinforcement Learning (RL) has shown remarkable success in enhancing the reasoning capabilities of Large Language Models (LLMs). Process-Supervised RL (PSRL) has emerged as a more…

cs.LG2024

RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors

Fengshuo Bai, Runze Liu, Yali Du +2

Evaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into sp…

cs.LG2024

PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation

Runze Liu, Yali Du, Fengshuo Bai +2

In preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot…

cs.LG2026

Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy

Jiashun Liu, Runze Liu, Xu Wan +3

Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses. Recent…

cs.AR2022

Eventor: An Efficient Event-Based Monocular Multi-View Stereo Accelerator on FPGA Platform

Mingjun Li, Jianlei Yang, Yingjie Qi +6

Event cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a tec…

cs.CV2025

A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding

Yida Wang, Taiting Lu, Runze Liu +12

Printed-Circuit-board (PCB) footprint geometry labeling of integrated circuits (IC) is essential in defining the physical interface between components and the PCB layout, requiring…

cs.CV2025

DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment

Junjie Gao, Runze Liu, Yingzhe Peng +4

Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide…

cond-mat.mtrl-sci2022

Calibrating DFT formation enthalpy calculations by multi-fidelity machine learning

Sheng Gong, Shuo Wang, Tian Xie +3

Machine learning materials properties measured by experiments is valuable yet difficult due to the limited amount of experimental data. In this work, we use a multi-fidelity random…

eess.SP2025

Chip-Surface Based Visual Authentication for Integrated Circuits

Runze Liu, Prasun Datta, Anirudh Nakra +2

The rapid development of the semiconductor industry and the ubiquity of electronic devices have led to a significant increase in the counterfeiting of integrated circuits (ICs). Th…

cond-mat.mtrl-sci2025

Materials Graph Library (MatGL), an open-source graph deep learning library for materials science and chemistry

Tsz Wai Ko, Bowen Deng, Marcel Nassar +7

Graph deep learning models, which incorporate a natural inductive bias for a collection of atoms, are of immense interest in materials science and chemistry. Here, we introduce the…