papers

Publications (33)

cs.CV2025

Voxel-based Point Cloud Geometry Compression with Space-to-Channel Context

Bojun Liu, Yangzhi Ma, Ao Luo +2

Voxel-based methods are among the most efficient for point cloud geometry compression, particularly with dense point clouds. However, they face limitations due to a restricted rece…

cs.CV2022

Learning Optical Flow with Adaptive Graph Reasoning

Ao Luo, Fan Yang, Kunming Luo +3

Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques l…

cs.CV2021

ASFlow: Unsupervised Optical Flow Learning with Adaptive Pyramid Sampling

Kunming Luo, Ao Luo, Chuan Wang +2

We present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propo…

cs.LG2026

Tyan-WP: A Wind Power Foundation Model for Ultra-Short-Term Probabilistic Forecasting

Jiahui Huang, Ao Luo, Lei Liu +6

Global wind power capacity, especially in China, is booming, with new farms spanning diverse terrains and climates. The industry urgently needs accurate wind power foundation model…

cs.CV2026

Off-the-shelf Vision Models Benefit Image Manipulation Localization

Zhengxuan Zhang, Keji Song, Junmin Hu +2

Image manipulation localization (IML) and general vision tasks are typically treated as two separate research directions due to the fundamental differences between manipulation-spe…

cs.CV2025

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model

Shijian Wang, Linxin Song, Jieyu Zhang +9

Current multimodal language model (MLM) training approaches overlook the influence of instruction templates. Previous research deals with this problem by leveraging hand-crafted or…

cs.CV2025

MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification

Jianwei Zhao, Xin Li, Fan Yang +5

Whole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance durin…

hep-ph2022

Enhancement of baryon-to-meson ratios around jets as a signature of medium response

Ao Luo, Ya-Xian Mao, Guang-You Qin +2

We present a unique signal of jet-induced medium excitations: the enhancement of baryon-to-meson ratios around the quenched jets.To illustrate this, we study jet-particle correlati…

cs.CV2024

SCP: Spherical-Coordinate-based Learned Point Cloud Compression

Ao Luo, Linxin Song, Keisuke Nonaka +4

In recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, the spinning LiDAR point cloud, is generated by spinning LiDAR…

cs.CV2024

LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models

Hai Jiang, Ao Luo, Xiaohong Liu +2

In this paper, we propose a diffusion-based unsupervised framework that incorporates physically explainable Retinex theory with diffusion models for low-light image enhancement, na…

eess.IV2022

Memory-Efficient Learned Image Compression with Pruned Hyperprior Module

Ao Luo, Heming Sun, Jinming Liu +1

Learned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. H…

cs.CL2023

Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training

Yan Zeng, Wangchunshu Zhou, Ao Luo +2

In this paper, we introduce Cross-View Language Modeling, a simple and effective pre-training framework that unifies cross-lingual and cross-modal pre-training with shared architec…

cs.LG2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

Kunyu Peng, Di Wen, Kailun Yang +6

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories ar…

cs.CL2021

WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Xingyu Chen, Zihan Zhao, Lu Chen +5

Web search is an essential way for humans to obtain information, but it's still a great challenge for machines to understand the contents of web pages. In this paper, we introduce…

cs.CV2026

DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesis

Xinglong Luo, Ao Luo, Zhengning Wang +5

Image alignment is a fundamental task in computer vision with broad applications. Existing methods predominantly employ optical flow-based image warping. However, this technique is…

hep-ph2022

Jet shape and redistribution of the lost energy from jets in Pb+Pb collisions at the LHC in a multiphase transport model

Ao Luo, Ya-Xian Mao, Guang-You Qin +2

Jet-medium interaction involves two important effects: jet energy loss and medium response. The search for jet-induced medium excitations is one of the hot topics in jet quenching…

cs.CV2025

Learning Efficient Meshflow and Optical Flow from Event Cameras

Xinglong Luo, Ao Luo, Kunming Luo +4

In this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start,…

cs.CV2020

Fast Portrait Segmentation with Highly Light-weight Network

Yuezun Li, Ao Luo, Siwei Lyu

In this paper, we describe a fast and light-weight portrait segmentation method based on a new highly light-weight backbone (HLB) architecture. The core element of HLB is a bottlen…

cs.CV2020

Deep-VFX: Deep Action Recognition Driven VFX for Short Video

Ao Luo, Ning Xie, Zhijia Tao +1

Human motion is a key function to communicate information. In the application, short-form mobile video is so popular all over the world such as Tik Tok. The users would like to add…

cs.CV2024

FastForensics: Efficient Two-Stream Design for Real-Time Image Manipulation Detection

Yangxiang Zhang, Yuezun Li, Ao Luo +2

With the rise in popularity of portable devices, the spread of falsified media on social platforms has become rampant. This necessitates the timely identification of authentic cont…

cs.CL2025

Adaptive In-conversation Team Building for Language Model Agents

Linxin Song, Jiale Liu, Jieyu Zhang +5

Leveraging multiple large language model (LLM) agents has shown to be a promising approach for tackling complex tasks, while the effective design of multiple agents for a particula…

cs.CV2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

Di Wen, Lei Qi, Kunyu Peng +9

Despite substantial progress in video understanding, most existing datasets are limited to Earth's gravitational conditions. However, microgravity alters human motion, interactions…

cs.CV2023

Low-Light Image Enhancement with Wavelet-based Diffusion Models

Hai Jiang, Ao Luo, Songchen Han +2

Diffusion models have achieved promising results in image restoration tasks, yet suffer from time-consuming, excessive computational resource consumption, and unstable restoration.…

hep-ph2018

Overall momentum balance and redistribution of the lost energy in asymmetric dijet events in 2.76~ATeV Pb-Pb collisions with a multi-phase transport model

Zhan Gao, Ao Luo, Guo-Liang Ma +2

The overall transverse momentum balance and the redistribution of the lost energy from hard jets for asymmetric dijet events in PbPb collisions at 2.76~ATeV at the LHC is studied w…

cs.CV2023

GAFlow: Incorporating Gaussian Attention into Optical Flow

Ao Luo, Fan Yang, Xin Li +4

Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving con…

cs.CV2022

RealFlow: EM-based Realistic Optical Flow Dataset Generation from Videos

Yunhui Han, Kunming Luo, Ao Luo +4

Obtaining the ground truth labels from a video is challenging since the manual annotation of pixel-wise flow labels is prohibitively expensive and laborious. Besides, existing appr…

nucl-th2024

Probing medium response via strangeness enhancement around quenched jets

Ao Luo, Shanshan Cao, Guang-You Qin

Jet-induced medium excitation is a crucial part of jet interactions with the quark-gluon plasma (QGP) in relativistic heavy-ion collisions, and has recently been confirmed by exper…

cs.CV2020

Cascade Graph Neural Networks for RGB-D Salient Object Detection

Ao Luo, Xin Li, Fan Yang +3

In this paper, we study the problem of salient object detection (SOD) for RGB-D images using both color and depth information.A major technical challenge in performing salient obje…

cs.CV2024

FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection

Jianwei Zhao, Xin Li, Fan Yang +4

Detecting objects seamlessly blended into their surroundings represents a complex task for both human cognitive capabilities and advanced artificial intelligence algorithms. Curren…

cs.CL2024

Better Explain Transformers by Illuminating Important Information

Linxin Song, Yan Cui, Ao Luo +2

Transformer-based models excel in various natural language processing (NLP) tasks, attracting countless efforts to explain their inner workings. Prior methods explain Transformers…

cs.CV2024

RecDiffusion: Rectangling for Image Stitching with Diffusion Models

Tianhao Zhou, Haipeng Li, Ziyi Wang +5

Image stitching from different captures often results in non-rectangular boundaries, which is often considered unappealing. To solve non-rectangular boundaries, current solutions i…

cs.CV2023

Learning Optical Flow from Event Camera with Rendered Dataset

Xinglong Luo, Kunming Luo, Ao Luo +3

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels…

cs.CV2020

Hybrid Graph Neural Networks for Crowd Counting

Ao Luo, Fan Yang, Xin Li +4

Crowd counting is an important yet challenging task due to the large scale and density variation. Recent investigations have shown that distilling rich relations among multi-scale…