Publications (40)
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
Jiakang Wang, Runze Liu, Qingpeng Cai +7
Reinforcement learning (RL) has shown great promise in large language models (LLMs) post-training, which typically rely on token-level clipping to maintain stability during optimiz…
Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions
Bingxu Liu, Jiashun Liu, Johan Obando-Ceron +5
While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stat…
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff
Runze Liu, Jiashun Liu, Xu Wan +2
Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT is expected to provide a usefu…
A Survey of Reinforcement Learning for Large Reasoning Models
Kaiyan Zhang, Yuxin Zuo, Bingxiang He +36
In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontie…
ReviewRL: Towards Automated Scientific Review with RL
Sihang Zeng, Kai Tian, Kaiyan Zhang +9
Peer review is essential for scientific progress but faces growing challenges due to increasing submission volumes and reviewer fatigue. Existing automated review approaches strugg…
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
Jiakang Wang, Runze Liu, Fuzheng Zhang +3
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training method for improving the reasoning abilities of Large Language Models (LLMs). However, e…
A hybrid single quantum dot coupled cavity on a CMOS-compatible SiC photonic chip for Purcell-enhanced deterministic single-photon emission
Yifan Zhu, Runze Liu, Ailun Yi +10
The ability to control nonclassical light emission from a single quantum emitter by an integrated cavity may unleash new perspectives for integrated photonic quantum applications.…
OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing
Taiting Lu, Kaiyuan Lin, Ziwei Dong +18
Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However, their ability to reason abou…
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Ye Lu, Yihan Yan, Zhaoyang Zhang +4
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens s…
Individualized Deepfake Detection Exploiting Traces Due to Double Neural-Network Operations
Mushfiqur Rahman, Runze Liu, Chau-Wai Wong +1
In today's digital landscape, journalists urgently require tools to verify the authenticity of facial images and videos depicting specific public figures before incorporating them…
RoeNets: Predicting Discontinuity of Hyperbolic Systems from Continuous Data
Shiying Xiong, Xingzhe He, Yunjin Tong +2
We introduce Roe Neural Networks (RoeNets) that can predict the discontinuity of the hyperbolic conservation laws (HCLs) based on short-term discontinuous and even continuous train…
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
Shengjie Sun, Jiafei Lyu, Runze Liu +4
Offline imitation learning (offline IL) enables training effective policies without requiring explicit reward annotations. Recent approaches attempt to estimate rewards for unlabel…
SEABO: A Simple Search-Based Method for Offline Imitation Learning
Jiafei Lyu, Xiaoteng Ma, Le Wan +3
Offline reinforcement learning (RL) has attracted much attention due to its ability in learning from static offline datasets and eliminating the need of interacting with the enviro…
OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction
Taiting Lu, Runze Liu, Ziwei Dong +18
Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the…
Bohdi: Heterogeneous LLM Fusion with Automatic Data Exploration
Junqi Gao, Zhichang Guo, Dazhi Zhang +5
Heterogeneous Large Language Model (LLM) fusion integrates the strengths of multiple source LLMs with different architectures into a target LLM with low computational overhead. Whi…
Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and Memorization
Xiaotao Jia, Jianlei Yang, Runze Liu +3
Bayesian method is capable of capturing real world uncertainties/incompleteness and properly addressing the over-fitting issue faced by deep neural networks. In recent years, Bayes…
A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning
Shengjie Sun, Runze Liu, Jiafei Lyu +3
Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often…
eSLAM: An Energy-Efficient Accelerator for Real-Time ORB-SLAM on FPGA Platform
Runze Liu, Jianlei Yang, Yiran Chen +1
Simultaneous Localization and Mapping (SLAM) is a critical task for autonomous navigation. However, due to the computational complexity of SLAM algorithms, it is very difficult to…
VLP: Vision-Language Preference Learning for Embodied Manipulation
Runze Liu, Chenjia Bai, Jiafei Lyu +3
Reward engineering is one of the key challenges in Reinforcement Learning (RL). Preference-based RL effectively addresses this issue by learning from human feedback. However, it is…
Unsupervised Monocular Depth Estimation Based on Hierarchical Feature-Guided Diffusion
Runze Liu, Dongchen Zhu, Guanghui Zhang +5
Unsupervised monocular depth estimation has received widespread attention because of its capability to train without ground truth. In real-world scenarios, the images may be blurry…
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Jian Zhao, Runze Liu, Kaiyan Zhang +8
Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However…
On Microstructure Estimation Using Flatbed Scanners for Paper Surface Based Authentication
Runze Liu, Chau-Wai Wong
Paper surfaces under the microscopic view are observed to be formed by intertwisted wood fibers. Such structures of paper surfaces are unique from one location to another and are a…
OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout
Taiting Lu, Kaiyuan Lin, Mingjia Wang +12
Recent large language models (LLMs) have demonstrated remarkable progress in 3D spatial reasoning, spatial grounding, and fine-grained geometric understanding. However, their abili…
A Fast, Accurate, and Reactive Equivariant Foundation Potential
Tsz Wai Ko, Runze Liu, Adesh Rohan Mishra +3
Electrostatics govern charge transfer and reactivity in materials. Yet, most foundation potentials (FPs) either do not explicitly model such interactions or pay a prohibitive scali…
A Foundational Potential Energy Surface Dataset for Materials
Aaron D. Kaplan, Runze Liu, Ji Qi +6
Accurate potential energy surface (PES) descriptions are essential for atomistic simulations of materials. Universal machine learning interatomic potentials (UMLIPs) offer…
Temporal Difference Learning with Constrained Initial Representations
Jiafei Lyu, Jingwen Yang, Zhongjian Qiao +5
Recently, there have been numerous attempts to enhance the sample efficiency of off-policy reinforcement learning (RL) agents when interacting with the environment, including archi…
Relayed-loop optical scan amplification
Harishankar Jayakumar, Christopher Warkentin, Deano Farinella +4
Many modern sensing, processing, and fabrication technologies depend upon the sequential scanning of laser light. Due to inertial, thermal, and electrical limitations, the speed of…
Modeling the Nonsmoothness of Modern Neural Networks
Runze Liu, Chau-Wai Wong, Huaiyu Dai
Modern neural networks have been successful in many regression-based tasks such as face recognition, facial landmark detection, and image generation. In this work, we investigate a…
SPINBIS: Spintronics based Bayesian Inference System with Stochastic Computing
Xiaotao Jia, Jianlei Yang, Pengcheng Dai +3
Bayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-in…
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Runze Liu, Junqi Gao, Jian Zhao +5
Test-Time Scaling (TTS) is an important method for improving the performance of Large Language Models (LLMs) by using additional computation during the inference phase. However, cu…
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
Runze Liu, Jiakang Wang, Yuling Shi +11
Reinforcement Learning (RL) has shown remarkable success in enhancing the reasoning capabilities of Large Language Models (LLMs). Process-Supervised RL (PSRL) has emerged as a more…
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
Fengshuo Bai, Runze Liu, Yali Du +2
Evaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into sp…
PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation
Runze Liu, Yali Du, Fengshuo Bai +2
In preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot…
Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy
Jiashun Liu, Runze Liu, Xu Wan +3
Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses. Recent…
Eventor: An Efficient Event-Based Monocular Multi-View Stereo Accelerator on FPGA Platform
Mingjun Li, Jianlei Yang, Yingjie Qi +6
Event cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a tec…
A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
Yida Wang, Taiting Lu, Runze Liu +12
Printed-Circuit-board (PCB) footprint geometry labeling of integrated circuits (IC) is essential in defining the physical interface between components and the PCB layout, requiring…
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
Junjie Gao, Runze Liu, Yingzhe Peng +4
Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide…
Calibrating DFT formation enthalpy calculations by multi-fidelity machine learning
Sheng Gong, Shuo Wang, Tian Xie +3
Machine learning materials properties measured by experiments is valuable yet difficult due to the limited amount of experimental data. In this work, we use a multi-fidelity random…
Chip-Surface Based Visual Authentication for Integrated Circuits
Runze Liu, Prasun Datta, Anirudh Nakra +2
The rapid development of the semiconductor industry and the ubiquity of electronic devices have led to a significant increase in the counterfeiting of integrated circuits (ICs). Th…
Materials Graph Library (MatGL), an open-source graph deep learning library for materials science and chemistry
Tsz Wai Ko, Bowen Deng, Marcel Nassar +7
Graph deep learning models, which incorporate a natural inductive bias for a collection of atoms, are of immense interest in materials science and chemistry. Here, we introduce the…