papers

Publications (56)

eess.IV2021

Real-Time Quantized Image Super-Resolution on Mobile NPUs, Mobile AI 2021 Challenge: Report

Andrey Ignatov, Radu Timofte, Maurizio Denna +20

Image super-resolution is one of the most popular computer vision problems with many important applications to mobile devices. While many solutions have been proposed for this task…

math.NA2026

Parametric Probabilistic Manifold Decomposition for Nonlinear Model Reduction

Jiaming Guo, Dunhui Xiao

Probabilistic Manifold Decomposition (PMD)\cite{doi:10.1137/25M1738863}, developed in our earlier work, provides a nonlinear model reduction by embedding high-dimensional dynamics…

cs.CV2024

Real-Time 4K Super-Resolution of Compressed AVIF Images. AIS 2024 Challenge Survey

Marcos V. Conde, Zhijun Lei, Wen Li +72

This paper introduces a novel benchmark as part of the AIS 2024 Real-Time Image Super-Resolution (RTSR) Challenge, which aims to upscale compressed images from 540p to 4K resolutio…

cs.DC2025

QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation

Xinguo Zhu, Shaohui Peng, Jiaming Guo +10

Developing high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While LLMs…

cs.CV2026

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

Tao Yu, Haopeng Jin, Hao Wang +18

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Ope…

hep-ph2025

Line Operators in the Left-Right Symmetric Model

Jiaming Guo, Rui Yang, Xun Xue

In this paper, we studied line operators in the Left-Right Symmetric Model. The gauge group of Left-Right Symmetric Electroweak Model is ${G} = SU(3) \times SU(2)_{L} \times SU(2)_…

cs.LG2025

Efficient Diffusion Planning with Temporal Diffusion

Jiaming Guo, Rui Zhang, Zerun Li +7

Diffusion planning is a promising method for learning high-performance policies from offline data. To avoid the impact of discrepancies between planning and reality on performance,…

cs.LG2025

Policy Constraint by Only Support Constraint for Offline Reinforcement Learning

Yunkai Gao, Jiaming Guo, Fan Wu +1

Offline reinforcement learning (RL) aims to optimize a policy by using pre-collected datasets, to maximize cumulative rewards. However, offline reinforcement learning suffers chall…

cs.SE2026

QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization

Changxin Ke, Rui Zhang, Jiaming Guo +10

Large Language Models (LLMs) achieve strong program repair performance but often suffer from over-editing, where excessive modifications overwrite correct code and hinder bug local…

cs.CV2025

World-Consistent Data Generation for Vision-and-Language Navigation

Yu Zhong, Rui Zhang, Zihao Zhang +9

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main…

cs.CL2024

Assessing and Understanding Creativity in Large Language Models

Yunpu Zhao, Rui Zhang, Wenyi Li +10

In the field of natural language processing, the rapid development of large language model (LLM) has attracted more and more attention. LLMs have shown a high level of creativity i…

cs.AR2025

QiMeng: Fully Automated Hardware and Software Design for Processor Chip

Rui Zhang, Yuanbo Wen, Shuyao Cheng +17

Processor chip design technology serves as a key frontier driving breakthroughs in computer science and related fields. With the rapid advancement of information technology, conven…

eess.IV2024

Efficient Real-world Image Super-Resolution Via Adaptive Directional Gradient Convolution

Long Peng, Yang Cao, Renjing Pei +5

Real-SR endeavors to produce high-resolution images with rich details while mitigating the impact of multiple degradation factors. Although existing methods have achieved impressiv…

cs.LG2022

Causality-driven Hierarchical Structure Discovery for Reinforcement Learning

Shaohui Peng, Xing Hu, Rui Zhang +9

Hierarchical reinforcement learning (HRL) effectively improves agents' exploration efficiency on tasks with sparse reward, with the guide of high-quality hierarchical structures (e…

math.NA2025

Nonlinear Model Reduction by Probabilistic Manifold Decomposition

Jiaming Guo, Dunhui Xiao

This paper presents a novel non-linear model reduction method: Probabilistic Manifold Decomposition (PMD), which provides a powerful framework for constructing non-intrusive reduce…

cs.CV2026

The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Jiatong Li, Zheng Chen, Kai Liu +91

This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the resulting outcomes. The challenge…

cs.LG2021

Hindsight Value Function for Variance Reduction in Stochastic Dynamic Environment

Jiaming Guo, Rui Zhang, Xishan Zhang +6

Policy gradient methods are appealing in deep reinforcement learning but suffer from high variance of gradient estimate. To reduce the variance, the state value function is applied…

cs.CV2022

A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution

Hengsheng Zhang, Xueyi Zou, Jiaming Guo +3

Online processing of compressed videos to increase their resolutions attracts increasing and broad attention. Video Super-Resolution (VSR) using recurrent neural network architectu…

cs.AI2024

Luban: Building Open-Ended Creative Agents via Autonomous Embodied Verification

Yuxuan Guo, Shaohui Peng, Jiaming Guo +15

Building open agents has always been the ultimate goal in AI research, and creative agents are the more enticing. Existing LLM agents excel at long-horizon tasks with well-defined…

cs.LG2019

Predicting Alzheimer's Disease by Hierarchical Graph Convolution from Positron Emission Tomography Imaging

Jiaming Guo, Wei Qiu, Xiang Li +3

Imaging-based early diagnosis of Alzheimer Disease (AD) has become an effective approach, especially by using nuclear medicine imaging techniques such as Positron Emission Topograp…

cs.CV2026

ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorization

Bingchen Li, Zhixin Wang, Fan Li +5

Old photos preserve invaluable historical memories, making their restoration and colorization highly desirable. While existing restoration models can address some degradation issue…

cs.CL2024

Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding

Lei Huang, Jiaming Guo, Guanhua He +5

Generating long-term texts such as novels using artificial intelligence has always been a challenge. A common approach is to use large language models (LLMs) to construct a hierarc…

cs.CV2019

Unsupervised Bi-directional Flow-based Video Generation from one Snapshot

Lu Sheng, Junting Pan, Jiaming Guo +3

Imagining multiple consecutive frames given one single snapshot is challenging, since it is difficult to simultaneously predict diverse motions from a single image and faithfully g…

cs.RO2026

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation

Yutai Li, Shaohui Peng, Jiaming Guo +8

Vision-Language-Action (VLA) models offer a promising paradigm for generalist robotic policies, yet their adaptation is hindered by data inefficiency and poor generalization. We ar…

cs.AI2026

Discrete Diffusion for Complex and Congested Multi-Agent Path Finding with Sparse Social Attention

Yuanzhe Wang, Tian Zhi, Zihang Wei +8

Multi-Agent Path Finding (MAPF) is a coordination problem that requires computing globally consistent, collision-free trajectories from individual start positions to assigned goal…

cs.LG2026

QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design

Lei Huang, Rui Zhang, Jiaming Guo +9

Large language models (LLMs) have shown promising capabilities in hardware description language (HDL) generation. However, existing approaches often rely on free-form natural langu…

cs.LG2023

Efficient Symbolic Policy Learning with Differentiable Symbolic Expression

Jiaming Guo, Rui Zhang, Shaohui Peng +8

Deep reinforcement learning (DRL) has led to a wide range of advances in sequential decision-making tasks. However, the complexity of neural network policies makes it difficult to…

cs.LG2023

Contrastive Modules with Temporal Attention for Multi-Task Reinforcement Learning

Siming Lan, Rui Zhang, Qi Yi +10

In the field of multi-task reinforcement learning, the modular principle, which involves specializing functionalities into different modules and combining them appropriately, has b…

cs.LG2023

Conceptual Reinforcement Learning for Language-Conditioned Tasks

Shaohui Peng, Xing Hu, Rui Zhang +7

Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recentl…

cs.AI2025

QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code

Hainan Fang, Yuanbo Wen, Jun Bi +8

Compilers, while essential, are notoriously complex systems that demand prohibitively expensive human expertise to develop and maintain. The recent advancements in Large Language M…

cs.SE2025

QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation

Changxin Ke, Rui Zhang, Shuo Wang +11

The rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel prog…

cs.LG2026

Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training

Xue Gong, Qi Yi, Ziyuan Nan +8

Training Large Language Models (LLMs) for reasoning tasks is increasingly driven by Reinforcement Learning with Verifiable Rewards (RLVR), where Proximal Policy Optimization (PPO)…

eess.IV2024

Unveiling Hidden Details: A RAW Data-Enhanced Paradigm for Real-World Super-Resolution

Long Peng, Wenbo Li, Jiaming Guo +7

Real-world image super-resolution (Real SR) aims to generate high-fidelity, detail-rich high-resolution (HR) images from low-resolution (LR) counterparts. Existing Real SR methods…

cs.CV2025

Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation

Yunpu Zhao, Rui Zhang, Junbin Xiao +5

Vision-language models (VLMs) excel in various multimodal tasks but frequently suffer from poor calibration, resulting in misalignment between their verbalized confidence and respo…

cs.LG2022

Object-Category Aware Reinforcement Learning

Qi Yi, Rui Zhang, Shaohui Peng +6

Object-oriented reinforcement learning (OORL) is a promising way to improve the sample efficiency and generalization ability over standard RL. Recent works that try to solve OORL t…

cs.CV2024

Prompt-based Visual Alignment for Zero-shot Policy Transfer

Haihan Gao, Rui Zhang, Qi Yi +13

Overfitting in RL has become one of the main obstacles to applications in reinforcement learning(RL). Existing methods do not provide explicit semantic constrain for the feature ex…

cs.LG2020

DWM: A Decomposable Winograd Method for Convolution Acceleration

Di Huang, Xishan Zhang, Rui Zhang +9

Winograd's minimal filtering algorithm has been widely used in Convolutional Neural Networks (CNNs) to reduce the number of multiplications for faster processing. However, it is on…

cs.LG2026

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

Enshuai Zhou, Yifan Hao, Chao Wang +7

Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentall…

physics.bio-ph2019

Three-Dimensional Microscopy by Milling with Ultraviolet Excitation

Jiaming Guo, Camille Artur, Jason L. Eriksen +1

Analysis of three-dimensional biological samples is critical to understanding tissue function and the mechanisms of disease. Many chronic conditions, like neurodegenerative disease…

cs.LG2020

Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers

Xishan Zhang, Shaoli Liu, Rui Zhang +8

Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers. Recent emerged quantization technique has been applied to inference of deep neur…

eess.IV2023

AsConvSR: Fast and Lightweight Super-Resolution Network with Assembled Convolutions

Jiaming Guo, Xueyi Zou, Yuyi Chen +4

In recent years, videos and images in 720p (HD), 1080p (FHD) and 4K (UHD) resolution have become more popular for display devices such as TVs, mobile phones and VR. However, these…

cs.CL2023

Self-driven Grounding: Large Language Model Agents with Automatical Language-aligned Skill Learning

Shaohui Peng, Xing Hu, Qi Yi +9

Large language models (LLMs) show their powerful automatic reasoning and planning capability with a wealth of semantic knowledge about the human world. However, the grounding probl…

cs.LG2023

Online Prototype Alignment for Few-shot Policy Transfer

Qi Yi, Rui Zhang, Shaohui Peng +10

Domain adaptation in reinforcement learning (RL) mainly deals with the changes of observation when transferring the policy to a new environment. Many traditional approaches of doma…

cs.CV2026

GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting

Mingyu Shi, Xin Di, Long Peng +8

Continuous Spatio-Temporal Video Super-Resolution (C-STVSR) aims to simultaneously enhance the spatial resolution and frame rate of videos by arbitrary scale factors, offering grea…

cs.LG2025

QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation

Yang Zhang, Rui Zhang, Jiaming Guo +10

The remarkable progress of Large Language Models (LLMs) presents promising opportunities for Verilog code generation which is significantly important for automated circuit design.…

cs.CL2025

QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach

Shouyang Dong, Yuanbo Wen, Jun Bi +10

Heterogeneous deep learning systems (DLS) such as GPUs and ASICs have been widely deployed in industrial data centers, which requires to develop multiple low-level tensor programs…

cs.AI2023

Emergent Communication for Rules Reasoning

Yuxuan Guo, Yifan Hao, Rui Zhang +14

Research on emergent communication between deep-learning-based agents has received extensive attention due to its inspiration for linguistics and artificial intelligence. However,…

cs.CV2025

Test-Time Preference Optimization for Image Restoration

Bingchen Li, Xin Li, Jiaqi Xu +4

Image restoration (IR) models are typically trained to recover high-quality images using L1 or LPIPS loss. To handle diverse unknown degradations, zero-shot IR methods have also be…

hep-ph2025

The Line operators in the G2HDM model

Rui Yang, Jiaming Guo, Linghai Li +2

We investigate the global structure of the Gauged Two-Higgs-Doublet Model (G2HDM), a framework that extends the Standard Model by introducing a dark sector governed by the gauge sy…

cs.CV2024

The Ninth NTIRE 2024 Efficient Super-Resolution Challenge Report

Bin Ren, Yawei Li, Nancy Mehta +129

This paper provides a comprehensive review of the NTIRE 2024 challenge, focusing on efficient single-image super-resolution (ESR) solutions and their outcomes. The task of this cha…

cs.RO2026

CoInfra: A Large-Scale Cooperative Infrastructure Perception System and Dataset for Vehicle-Infrastructure Cooperation in Adverse Weather

Minghao Ning, Yufeng Yang, Keqi Shu +10

Vehicle-infrastructure (V2I) cooperative perception can substantially extend the range, coverage, and robustness of autonomous driving systems beyond the limits of onboard-only sen…

cs.LG2023

Context Shift Reduction for Offline Meta-Reinforcement Learning

Yunkai Gao, Rui Zhang, Jiaming Guo +10

Offline meta-reinforcement learning (OMRL) utilizes pre-collected offline datasets to enhance the agent's generalization ability on unseen tasks. However, the context shift problem…

eess.IV2021

Real-Time Video Super-Resolution on Smartphones with Deep Learning, Mobile AI 2021 Challenge: Report

Andrey Ignatov, Andres Romero, Heewon Kim +28

Video super-resolution has recently become one of the most important mobile-related problems due to the rise of video communication and streaming services. While many solutions hav…

eess.IV2019

Multi-label Detection and Classification of Red Blood Cells in Microscopic Images

Wei Qiu, Jiaming Guo, Xiang Li +4

Cell detection and cell type classification from biomedical images play an important role for high-throughput imaging and various clinical application. While classification of sing…

cs.AI2025

Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation

Yu Zhong, Zihao Zhang, Rui Zhang +9

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of har…

cs.AI2025

Code Driven Planning with Domain-Adaptive Critic

Zikang Tian, Shaohui Peng, Du Huang +11

Large Language Models (LLMs) have been widely adopted as task planners for AI agents in sequential decision-making problems, leveraging their extensive world knowledge. However, th…