papers

Publications (43)

cs.CV2023

Multistage Spatial Context Models for Learned Image Compression

Fangzheng Lin, Heming Sun, Jinming Liu +1

Recent state-of-the-art Learned Image Compression methods feature spatial context models, achieving great rate-distortion improvements over hyperprior methods. However, the autoreg…

eess.IV2020

Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

Image compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fa…

eess.IV2024

LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression

Shimon Murai, Heming Sun, Jiro Katto

Supported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achi…

eess.IV2021

Learned Video Compression with Residual Prediction and Loop Filter

Chao Liu, Heming Sun, Jiro Katto +2

In this paper, we propose a learned video codec with a residual prediction network (RP-Net) and a feature-aided loop filter (LF-Net). For the RP-Net, we exploit the residual of pre…

cs.CV2024

Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs

Jinming Liu, Yuntao Wei, Junyan Lin +5

We present a new image compression paradigm to achieve ``intelligently coding for machine'' by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motiva…

eess.IV2022

Q-LIC: Quantizing Learned Image Compression with Channel Splitting

Heming Sun, Lu Yu, Jiro Katto

Learned image compression (LIC) has reached a comparable coding gain with traditional hand-crafted methods such as VVC intra. However, the large network complexity prohibits the us…

eess.IV2023

Learned Image Compression with Mixed Transformer-CNN Architectures

Jinming Liu, Heming Sun, Jiro Katto

Learned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existin…

cs.AI2026

LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models

Tiesunlong Shen, Rui Mao, Jin Wang +4

Aligning Large Language Models (LLMs) with human preferences is critical, yet traditional fine-tuning methods are computationally expensive and inflexible. While test-time alignmen…

cs.AR2024

Multi-diseases detection with memristive system on chip

Zihan Wang, Daniel W. Yang, Zerui Liu +5

This study presents the first implementation of multilayer neural networks on a memristor/CMOS integrated system on chip (SoC) to simultaneously detect multiple diseases. To overco…

eess.IV2019

Dual Learning-based Video Coding with Inception Dense Blocks

Chao Liu, Heming Sun, Junan Chen +5

In this paper, a dual learning-based method in intra coding is introduced for PCS Grand Challenge. This method is mainly composed of two parts: intra prediction and reconstruction…

cs.CV2024

SCP: Spherical-Coordinate-based Learned Point Cloud Compression

Ao Luo, Linxin Song, Keisuke Nonaka +4

In recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, the spinning LiDAR point cloud, is generated by spinning LiDAR…

cs.CV2018

Deep Convolutional AutoEncoder-based Lossy Image Compression

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

Image compression has been investigated as a fundamental research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and is g…

cs.CV2021

Learned Image Compression with Separate Hyperprior Decoders

Zhao Zan, Chao Liu, Heming Sun +2

Learned image compression techniques have achieved considerable development in recent years. In this paper, we find that the performance bottleneck lies in the use of a single hype…

eess.IV2020

End-to-end Learned Image Compression with Fixed Point Weight Quantization

Heming Sun, Zhengxue Cheng, Masaru Takeuchi +1

Learned image compression (LIC) has reached the traditional hand-crafted methods such as JPEG2000 and BPG in terms of the coding gain. However, the large model size of the network…

eess.IV2022

Memory-Efficient Learned Image Compression with Pruned Hyperprior Module

Ao Luo, Heming Sun, Jinming Liu +1

Learned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. H…

cs.CV2024

Lightweight Stochastic Video Prediction via Hybrid Warping

Kazuki Kotoyori, Shota Hirose, Heming Sun +1

Accurate video prediction by deep neural networks, especially for dynamic regions, is a challenging task in computer vision for critical applications such as autonomous driving, re…

cs.CV2022

ABCAS: Adaptive Bound Control of spectral norm as Automatic Stabilizer

Shota Hirose, Shiori Maki, Naoki Wada +2

Spectral Normalization is one of the best methods for stabilizing the training of Generative Adversarial Network. Spectral Normalization limits the gradient of discriminator betwee…

eess.IV2021

End-to-End Learned Image Compression with Quantized Weights and Activations

Heming Sun, Lu Yu, Jiro Katto

End-to-end Learned image compression (LIC) has reached the traditional hand-crafted methods such as BPG (HEVC intra) in terms of the coding gain. However, the large network size pr…

eess.IV2022

Learned Lossless Image Compression With Combined Autoregressive Models And Attention Modules

Ran Wang, Jinming Liu, Heming Sun +1

Lossless image compression is an essential research field in image compression. Recently, learning-based image compression methods achieved impressive performance compared with tra…

eess.IV2025

A Multi-Grid Implicit Neural Representation for Multi-View Videos

Qingyue Ling, Zhengxue Cheng, Donghui Feng +6

Multi-view videos are becoming widely used in different fields, but their high resolution and multi-camera shooting raise significant challenges for storage and transmission. In th…

cs.CV2023

Prompt-ICM: A Unified Framework towards Image Coding for Machines with Task-driven Prompts

Ruoyu Feng, Jinming Liu, Xin Jin +3

Image coding for machines (ICM) aims to compress images to support downstream AI analysis instead of human perception. For ICM, developing a unified codec to reduce information red…

eess.IV2024

Attack and Defense Analysis of Learned Image Compression

Tianyu Zhu, Heming Sun, Xiankui Xiong +4

Learned image compression (LIC) is becoming more and more popular these years with its high efficiency and outstanding compression quality. Still, the practicality against modified…

eess.IV2020

A Convolutional Neural Network-Based Low Complexity Filter

Chao Liu, Heming Sun, Jiro Katto +2

Convolutional Neural Network (CNN)-based filters have achieved significant performance in video artifacts reduction. However, the high complexity of existing methods makes it diffi…

eess.IV2019

Learning Image and Video Compression through Spatial-Temporal Energy Compaction

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

Compression has been an important research topic for many decades, to produce a significant impact on data transmission and storage. Recent advances have shown a great potential of…

eess.IV2020

Learned Lossless Image Compression with a HyperPrior and Discretized Gaussian Mixture Likelihoods

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

Lossless image compression is an important task in the field of multimedia communication. Traditional image codecs typically support lossless mode, such as WebP, JPEG2000, FLIF. Re…

eess.IV2020

A QP-adaptive Mechanism for CNN-based Filter in Video Coding

Chao Liu, Heming Sun, Jiro Katto +2

Convolutional neural network (CNN)-based filters have achieved great success in video coding. However, in most previous works, individual models are needed for each quantization pa…

eess.IV2020

Enhanced Intra Prediction for Video Coding by Using Multiple Neural Networks

Heming Sun, Zhengxue Cheng, Masaru Takeuchi +1

This paper enhances the intra prediction by using multiple neural network modes (NM). Each NM serves as an end-to-end mapping from the neighboring reference blocks to the current c…

eess.IV2026

Streamable Neural Video Compression: A Mixed Precision Approach for Cross-Platform Deployment

Kasidis Arunruangsirilert, Heming Sun, Jiro Katto

Neural Video Codecs (NVCs) offer unprecedented rate-distortion performance, making them highly attractive for bandwidth-constrained environments like 5G cellular networks and emerg…

cs.CV2025

Real-time Video Prediction With Fast Video Interpolation Model and Prediction Training

Shota Hirose, Kazuki Kotoyori, Kasidis Arunruangsirilert +3

Transmission latency significantly affects users' quality of experience in real-time interaction and actuation. As latency is principally inevitable, video prediction can be utiliz…

cs.DC2023

Recoil: Parallel rANS Decoding with Decoder-Adaptive Scalability

Fangzheng Lin, Kasidis Arunruangsirilert, Heming Sun +1

Entropy coding is essential to data compression, image and video coding, etc. The Range variant of Asymmetric Numeral Systems (rANS) is a modern entropy coder, featuring superior s…

cs.CV2026

Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion

Jinming Liu, Yuntao Wei, Junyan Lin +5

Learned image compression methods have shown impressive performance but are often highly specialized for either human perception or specific machine vision tasks. This specializati…

eess.IV2019

Deep Residual Learning for Image Compression

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

In this paper, we provide a detailed description on our approach designed for CVPR 2019 Workshop and Challenge on Learned Image Compression (CLIC). Our approach mainly consists of…

cs.CV2026

Dual-Constrained Diffusion Image Compression for Operational Rate-Distortion-Perception Optimization

Sanxin Jiang, Jiro Katto, Heming Sun

The rate-distortion-perception (RDP) trade-off extends classical rate--distortion theory by imposing a distributional constraint on reconstructions, providing a unified framework f…

cs.CV2022

Semantic Segmentation in Learned Compressed Domain

Jinming Liu, Heming Sun, Jiro Katto

Most machine vision tasks (e.g., semantic segmentation) are based on images encoded and decoded by image compression algorithms (e.g., JPEG). However, these decoded images in the p…

eess.IV2019

Perceptual Quality Study on Deep Learning based Image Compression

Zhengxue Cheng, Pinar Akyazi, Heming Sun +2

Recently deep learning based image compression has made rapid advances with promising results based on objective quality metrics. However, a rigorous subjective quality evaluation…

eess.IV2023

Accelerating Learnt Video Codecs with Gradient Decay and Layer-wise Distillation

Tianhao Peng, Ge Gao, Heming Sun +2

In recent years, end-to-end learnt video codecs have demonstrated their potential to compete with conventional coding algorithms in term of compression efficiency. However, most le…

eess.IV2020

Low Bitrate Image Compression with Discretized Gaussian Mixture Likelihoods

Zhengxue Cheng, Heming Sun, Jiro Katto

In this paper, we provide a detailed description on our submitted method Kattolab to Workshop and Challenge on Learned Image Compression (CLIC) 2020. Our method mainly incorporates…

eess.IV2018

Performance Comparison of Convolutional AutoEncoders, Generative Adversarial Networks and Super-Resolution for Image Compression

Zhengxue Cheng, Heming Sun, Masaru Takeuchi +1

Image compression has been investigated for many decades. Recently, deep learning approaches have achieved a great success in many computer vision tasks, and are gradually used in…

cs.AR2021

FPGA Based Accelerator for Neural Networks Computation with Flexible Pipelining

Qingyang Yi, Heming Sun, Masahiro Fujita

FPGA is appropriate for fix-point neural networks computing due to high power efficiency and configurability. However, its design must be intensively refined to achieve high perfor…

cs.CL2021

COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval

Xinliang Frederick Zhang, Heming Sun, Xiang Yue +2

We present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval. Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set…

eess.IV2024

Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization

Zhibo Chen, Heming Sun, Li Zhang +1

This paper provides a survey of the latest developments in visual signal coding and processing with generative models. Specifically, our focus is on presenting the advancement of g…

eess.IV2022

Streaming-capable High-performance Architecture of Learned Image Compression Codecs

Fangzheng Lin, Heming Sun, Jiro Katto

Learned image compression allows achieving state-of-the-art accuracy and compression ratios, but their relatively slow runtime performance limits their usage. While previous attemp…

eess.IV2021

Fully Neural Network Mode Based Intra Prediction of Variable Block Size

Heming Sun, Lu Yu, Jiro Katto

Intra prediction is an essential component in the image coding. This paper gives an intra prediction framework completely based on neural network modes (NM). Each NM can be regarde…