papers

Publications (24)

cs.CV2026

Frequency-aware Neural Representation for Videos

Jun Zhu, Xinfeng Zhang, Lv Tang +3

Implicit Neural Representations (INRs) have emerged as a promising paradigm for video compression. However, existing INR-based frameworks typically suffer from inherent spectral bi…

cs.CV2021

Highly Efficient Natural Image Matting

Yijie Zhong, Bo Li, Lv Tang +2

Over the last few years, deep learning based approaches have achieved outstanding improvements in natural image matting. However, there are still two drawbacks that impede the wide…

eess.IV2024

Releasing the Parameter Latency of Neural Representation for High-Efficiency Video Compression

Gai Zhang, Xinfeng Zhang, Lv Tang +3

For decades, video compression technology has been a prominent research area. Traditional hybrid video compression framework and end-to-end frameworks continue to explore various i…

cs.LG2026

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs

Lv Tang, Tianyi Zheng, Bo Li +1

Unified multimodal large language models (MLLMs) aim to unify image understanding and image generation within a single framework, where a shared visual tokenizer serves as the sole…

cs.CV2026

Visual Text Compression as Measure Transport

Lv Tang, Tianyi Zheng, Yang Liu +2

Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing --$20\t…

cs.CV2025

MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion

Jun Zhu, Xinfeng Zhang, Lv Tang +1

Implicit Neural representations (INRs) have emerged as a promising approach for video compression, and have achieved comparable performance to the state-of-the-art codecs such as H…

cs.CV2024

Evaluating SAM2's Role in Camouflaged Object Detection: From SAM to SAM2

Lv Tang, Bo Li

The Segment Anything Model (SAM), introduced by Meta AI Research as a generic object segmentation model, quickly garnered widespread attention and significantly influenced the acad…

cs.CV2024

Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection

Lv Tang, Peng-Tao Jiang, Zhihao Shen +3

In this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilit…

cs.CV2024

Zero-Shot Co-salient Object Detection Framework

Haoke Xiao, Lv Tang, Bo Li +2

Co-salient Object Detection (CoSOD) endeavors to replicate the human visual system's capacity to recognize common and salient objects within a collection of images. Despite recent…

cs.CV2020

CLASS: Cross-Level Attention and Supervision for Salient Objects Detection

Lv Tang, Bo Li

Salient object detection (SOD) is a fundamental computer vision task. Recently, with the revival of deep neural networks, SOD has made great progresses. However, there still exist…

cs.CV2021

Disentangled High Quality Salient Object Detection

Lv Tang, Bo Li, Shouhong Ding +1

Aiming at discovering and locating most distinctive objects from visual scenes, salient object detection (SOD) plays an essential role in various computer vision systems. Coming to…

cs.CV2025

UAR-NVC: A Unified AutoRegressive Framework for Memory-Efficient Neural Video Compression

Jia Wang, Xinfeng Zhang, Gai Zhang +3

Implicit Neural Representations (INRs) have demonstrated significant potential in video compression by representing videos as neural networks. However, as the number of frames incr…

cs.CV2026

FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry

Bo Li, Haoke Xiao, Lv Tang

Vision Mamba offers linear complexity for long visual sequences, yet its performance depends critically on how a two-dimensional patch grid is serialized into a one-dimensional sta…

cs.CV2025

CANeRV: Content Adaptive Neural Representation for Video Compression

Lv Tang, Jun Zhu, Xinfeng Zhang +3

Recent advances in video compression introduce implicit neural representation (INR) based methods, which effectively capture global dependencies and characteristics of entire video…

cs.CV2026

Video Compression with Hierarchical Temporal Neural Representation

Jun Zhu, Xinfeng Zhang, Lv Tang +3

Video compression has recently benefited from implicit neural representations (INRs), which model videos as continuous functions. INRs offer compact storage and flexible reconstruc…

cs.CV2022

Towards Stable Co-saliency Detection and Object Co-segmentation

Bo Li, Lv Tang, Senyun Kuang +2

In this paper, we present a novel model for simultaneous stable co-saliency detection (CoSOD) and object co-segmentation (CoSEG). To detect co-saliency (segmentation) accurately, t…

eess.IV2025

SANR: Scene-Aware Neural Representation for Light Field Image Compression with Rate-Distortion Optimization

Gai Zhang, Xinfeng Zhang, Lv Tang +3

Light field images capture multi-view scene information and play a crucial role in 3D scene reconstruction. However, their high-dimensional nature results in enormous data volumes,…

cs.CV2024

Towards Training-free Open-world Segmentation via Image Prompt Foundation Models

Lv Tang, Peng-Tao Jiang, Hao-Ke Xiao +1

The realm of computer vision has witnessed a paradigm shift with the advent of foundational models, mirroring the transformative influence of large language models in the domain of…

cs.AI2026

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

Changxin Huang, Lv Tang, Zhaohuan Zhan +5

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging.…

cs.CV2024

Scalable Visual State Space Model with Fractal Scanning

Lv Tang, HaoKe Xiao, Peng-Tao Jiang +3

Foundational models have significantly advanced in natural language processing (NLP) and computer vision (CV), with the Transformer architecture becoming a standard backbone. Howev…

cs.CV2024

ASAM: Boosting Segment Anything Model with Adversarial Tuning

Bo Li, Haoke Xiao, Lv Tang

In the evolving landscape of computer vision, foundation models have emerged as pivotal tools, exhibiting exceptional adaptability to a myriad of tasks. Among these, the Segment An…

cs.CV2023

Can SAM Segment Anything? When SAM Meets Camouflaged Object Detection

Lv Tang, Haoke Xiao, Bo Li

SAM is a segmentation model recently released by Meta AI Research and has been gaining attention quickly due to its impressive performance in generic object segmentation. However,…

cs.CV2022

CoSformer: Detecting Co-Salient Object with Transformers

Lv Tang, Bo Li

Co-Salient Object Detection (CoSOD) aims at simulating the human visual system to discover the common and salient objects from a group of relevant images. Recent methods typically…

cs.CV2023

Scene Matters: Model-based Deep Video Compression

Lv Tang, Xinfeng Zhang, Gai Zhang +1

Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal pred…