3 papers
cs.CV2026
Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
Tianyu Zhang, Dong Liu, Chang Wen Chen
Ultra-low bitrate image compression (below 0.05 bits per pixel) is increasingly critical for bandwidth-constrained and computation-limited encoding scenarios such as edge devices.…
cs.CV2025
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
Hongchen Wei, Zhihong Tan, Yaosi Hu +2
Large Multimodal Models (LMMs) have demonstrated exceptional performance in video captioning tasks, particularly for short videos. However, as the length of the video increases, ge…
cs.CV2024
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
Binyuan Huang, Yuqing Wen, Yucheng Zhao +9
Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data…