Publications (24)
Frequency-aware Neural Representation for Videos
Jun Zhu, Xinfeng Zhang, Lv Tang +3
Implicit Neural Representations (INRs) have emerged as a promising paradigm for video compression. However, existing INR-based frameworks typically suffer from inherent spectral bi…
Highly Efficient Natural Image Matting
Yijie Zhong, Bo Li, Lv Tang +2
Over the last few years, deep learning based approaches have achieved outstanding improvements in natural image matting. However, there are still two drawbacks that impede the wide…
Releasing the Parameter Latency of Neural Representation for High-Efficiency Video Compression
Gai Zhang, Xinfeng Zhang, Lv Tang +3
For decades, video compression technology has been a prominent research area. Traditional hybrid video compression framework and end-to-end frameworks continue to explore various i…
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
Lv Tang, Tianyi Zheng, Bo Li +1
Unified multimodal large language models (MLLMs) aim to unify image understanding and image generation within a single framework, where a shared visual tokenizer serves as the sole…
Visual Text Compression as Measure Transport
Lv Tang, Tianyi Zheng, Yang Liu +2
Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing --$20\t…
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
Jun Zhu, Xinfeng Zhang, Lv Tang +1
Implicit Neural representations (INRs) have emerged as a promising approach for video compression, and have achieved comparable performance to the state-of-the-art codecs such as H…
Evaluating SAM2's Role in Camouflaged Object Detection: From SAM to SAM2
Lv Tang, Bo Li
The Segment Anything Model (SAM), introduced by Meta AI Research as a generic object segmentation model, quickly garnered widespread attention and significantly influenced the acad…
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
Lv Tang, Peng-Tao Jiang, Zhihao Shen +3
In this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilit…
Zero-Shot Co-salient Object Detection Framework
Haoke Xiao, Lv Tang, Bo Li +2
Co-salient Object Detection (CoSOD) endeavors to replicate the human visual system's capacity to recognize common and salient objects within a collection of images. Despite recent…
CLASS: Cross-Level Attention and Supervision for Salient Objects Detection
Lv Tang, Bo Li
Salient object detection (SOD) is a fundamental computer vision task. Recently, with the revival of deep neural networks, SOD has made great progresses. However, there still exist…
Disentangled High Quality Salient Object Detection
Lv Tang, Bo Li, Shouhong Ding +1
Aiming at discovering and locating most distinctive objects from visual scenes, salient object detection (SOD) plays an essential role in various computer vision systems. Coming to…
UAR-NVC: A Unified AutoRegressive Framework for Memory-Efficient Neural Video Compression
Jia Wang, Xinfeng Zhang, Gai Zhang +3
Implicit Neural Representations (INRs) have demonstrated significant potential in video compression by representing videos as neural networks. However, as the number of frames incr…
FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
Bo Li, Haoke Xiao, Lv Tang
Vision Mamba offers linear complexity for long visual sequences, yet its performance depends critically on how a two-dimensional patch grid is serialized into a one-dimensional sta…
CANeRV: Content Adaptive Neural Representation for Video Compression
Lv Tang, Jun Zhu, Xinfeng Zhang +3
Recent advances in video compression introduce implicit neural representation (INR) based methods, which effectively capture global dependencies and characteristics of entire video…
Video Compression with Hierarchical Temporal Neural Representation
Jun Zhu, Xinfeng Zhang, Lv Tang +3
Video compression has recently benefited from implicit neural representations (INRs), which model videos as continuous functions. INRs offer compact storage and flexible reconstruc…
Towards Stable Co-saliency Detection and Object Co-segmentation
Bo Li, Lv Tang, Senyun Kuang +2
In this paper, we present a novel model for simultaneous stable co-saliency detection (CoSOD) and object co-segmentation (CoSEG). To detect co-saliency (segmentation) accurately, t…
SANR: Scene-Aware Neural Representation for Light Field Image Compression with Rate-Distortion Optimization
Gai Zhang, Xinfeng Zhang, Lv Tang +3
Light field images capture multi-view scene information and play a crucial role in 3D scene reconstruction. However, their high-dimensional nature results in enormous data volumes,…
Towards Training-free Open-world Segmentation via Image Prompt Foundation Models
Lv Tang, Peng-Tao Jiang, Hao-Ke Xiao +1
The realm of computer vision has witnessed a paradigm shift with the advent of foundational models, mirroring the transformative influence of large language models in the domain of…
UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
Changxin Huang, Lv Tang, Zhaohuan Zhan +5
Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging.…
Scalable Visual State Space Model with Fractal Scanning
Lv Tang, HaoKe Xiao, Peng-Tao Jiang +3
Foundational models have significantly advanced in natural language processing (NLP) and computer vision (CV), with the Transformer architecture becoming a standard backbone. Howev…
ASAM: Boosting Segment Anything Model with Adversarial Tuning
Bo Li, Haoke Xiao, Lv Tang
In the evolving landscape of computer vision, foundation models have emerged as pivotal tools, exhibiting exceptional adaptability to a myriad of tasks. Among these, the Segment An…
Can SAM Segment Anything? When SAM Meets Camouflaged Object Detection
Lv Tang, Haoke Xiao, Bo Li
SAM is a segmentation model recently released by Meta AI Research and has been gaining attention quickly due to its impressive performance in generic object segmentation. However,…
CoSformer: Detecting Co-Salient Object with Transformers
Lv Tang, Bo Li
Co-Salient Object Detection (CoSOD) aims at simulating the human visual system to discover the common and salient objects from a group of relevant images. Recent methods typically…
Scene Matters: Model-based Deep Video Compression
Lv Tang, Xinfeng Zhang, Gai Zhang +1
Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal pred…