1 citations · 1 across the 5 of their papers we have counts for
6 papers · 1 filter
Programmable World Model
Zheng-Hui Huang, Guixu Lin, Jiacheng Lin +8
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing prog…
Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution
Ren Wang, Yung-Yu Chuang
The perception-distortion trade-off poses a fundamental challenge in single-image super-resolution (SR). Although diffusion-based SR methods excel at generating perceptually realis…
Reflection Separation from a Single Image via Joint Latent Diffusion
Zheng-Hui Huang, Zhixiang Wang, Yu-Lun Liu +1
Single-image reflection separation is highly challenging under extreme conditions like glare or weak reflections. Existing methods often struggle to recover both layers in glare or…
Generative World Renderer
Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan +6
Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge thi…
Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation
Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang +6
This paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs wit…
2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
Cheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang +1
We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3…