computer vision

DCVC-MB: Neural B-Frame Video Compression using State Space Models

arXiv:2607.14305

summary

The paper introduces DCVC-MB, a neural video codec for low-delay B‑frame compression that uses state‑space models for bidirectional prediction and an entropy‑aware skipping mechanism to improve compression efficiency.

Abstract

In this paper we propose DCVC-Mamba (DCVC-MB), a neural video codec framework for B-frame coding. Our approach incorporates an IBP frame strategy for low-delay B-frame coding, a spatio-temporal fusion model based on state-space models for bidirectional temporal prediction, and an entropy-aware skipping mechanism that selectively omits coding certain latents to reduce entropy coding times. In addition to our model contributions we also implement two inference-time strategies that enhance compression performance. Experimental evaluation shows that DCVC-MB compares favorably to existing NVCs and traditional codecs. The method demonstrates BD-rate reductions of up to on average compared to prior neural video codecs, and improvements of up to and over the VTM-19.0-LDP and VTM-19.0-RA(Inter-GoP=16) benchmarks, respectively, contributing to advances in neural video compression.

Accepted to ICME 2026

Topics & keywords

#neural video compression#b-frame coding#state space models#entropy coding#low-delay videoDCVC-MBMambaIBP frame strategyspatio-temporal fusionentropy-aware skippingBD-rate
DCVC-MB: Neural B-Frame Video Compression using State Space Models · wovepaper