collaborators

11 papers

cs.LG2026

Training-Free Vector Quantization via Gaussian VAEs

Tongda Xu, Wendi Zheng, Jiajun He +4

Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images into discrete tokens. However, they are difficult to train due to discretization.…

cs.CV2026

Benchmarking and Enhancing VLM for Compressed Image Understanding

Zifu Zhang, Tongda Xu, Siqi Li +4

With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs has become increasingly importa…

cs.CV2026

SoLAR: Error-Resilient Streamable Long-Horizon Free-Viewpoint Video Reconstruction with Anchor Activation and Latent Recalibration

Haotian Zhang, Xu Mo, Yixin Yu +7

Free-Viewpoint Video (FVV) has emerged as a cornerstone of next-generation immersive media systems and attracted widespread attention. Previous methods primarily focus on short vid…

cs.CV2026

Making Reconstruction FID Predictive of Diffusion Generation FID

Tongda Xu, Mingwei He, Shady Abu-Hussein +6

It is well known that the reconstruction FID (rFID) of a VAE is poorly correlated with the generation FID (gFID) of a latent diffusion model. We propose interpolated FID (iFID), a…

cs.CV2026

Training-Free Image Editing with Visual Context Integration and Concept Alignment

Rui Song, Guo-Hua Wang, Qing-Guo Chen +6

In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visua…

cs.CV2026

RAWIC: Bit-Depth Adaptive Lossless Raw Image Compression

Chunhang Zheng, Tongda Xu, Mingli Xie +2

Raw images preserve linear sensor measurements and high bit-depth information crucial for advanced vision tasks and photography applications, yet their storage remains challenging…