activity
20182024
most citedBiFuse++: Self-supervised and Efficient Bi-projection Fusion for 360 Depth Estimation

33 citations · 82 across the 17 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CV2023

Transformer-based Image Compression with Variable Image Quality Objectives

Chia-Hao Kao, Yi-Hsin Chen, Cheng Chien +2

This paper presents a Transformer-based image compression system that allows for a variable image quality objective according to the user's preference. Optimizing a learned codec f…

cs.CV2023

Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where

Zhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang +2

While image data starts to enjoy the simple-but-effective self-supervised learning scheme built upon masking and self-reconstruction objective thanks to the introduction of tokeniz…

cs.CL2023

Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts

Zhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang +2

Text-to-image diffusion models, e.g. Stable Diffusion (SD), lately have shown remarkable ability in high-quality content generation, and become one of the representatives for the r…

eess.IV2023

TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception

Yi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao +3

This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Tra…

eess.IV2023

Transformer-based Variable-rate Image Compression with Region-of-interest Control

Chia-Hao Kao, Ying-Chieh Weng, Yi-Hsin Chen +2

This paper proposes a transformer-based learned image compression system. It is capable of achieving variable-rate compression with a single model while supporting the region-of-in…

cs.CV202310 cited

Multimodal Prompting with Missing Modalities for Visual Recognition

Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu +1

In this paper, we tackle two challenges in multimodal learning for visual recognition: 1) when missing-modality occurs either during training or testing in real-world situations; a…