12 papers
MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou +2
Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both re…
FlowCodec: One-Step Flow Prior for Generative Image Compression
Yinhuan Huang, Hao Cao, Pu chen +2
Diffusion-based image compression methods, leveraging powerful generative priors, have demonstrated remarkable perceptual quality at ultra-low bitrates. However, adapting modern ge…
ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding
Chengbin Liang, Wenqi Guo, Hao Cao +1
Neural speech codecs enable low-bitrate speech communication, yet at ultra-low bitrates (< 1000 bps) preserving perceptual quality and intelligibility is challenging. Existing desi…
Efficient Learned Image Compression without Entropy Coding
Hao Cao, Wenqi Guo, Zhijin Qin +1
Entropy coding is widely used in typical learned image compression (LIC) that converts latents into a compact bitstream. However, entropy coding is typically sequential and becomes…
ProGIC: Progressive and Lightweight Generative Image Compression with Residual Vector Quantization
Hao Cao, Chengbin Liang, Wenqi Guo +2
Recent advances in generative image compression (GIC) have delivered remarkable improvements in perceptual quality. However, many GICs rely on large-scale and rigid models, which s…
Position: Universal Aesthetic Alignment Narrows Artistic Expression
Wenqi Marshall Guo, Qingyun Qian, Khalad Hasan +1
Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when "anti-aesthetic" outputs are requested for artistic or cri…