3 papers
cs.CV2024
Bridging Compressed Image Latents and Multimodal Large Language Models
Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng +5
This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLM…
cs.CV2023
Transformer-based Image Compression with Variable Image Quality Objectives
Chia-Hao Kao, Yi-Hsin Chen, Cheng Chien +2
This paper presents a Transformer-based image compression system that allows for a variable image quality objective according to the user's preference. Optimizing a learned codec f…
eess.IV2023
TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception
Yi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao +3
This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Tra…