activity
20232026
collaborators

8 papers

cs.CV2026

CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction

Shiu-hong Kao, Chak Ho Huang, Huaiqian Liu +2

Existing works of reasoning segmentation often fall short in complex cases, particularly when addressing complicated queries and out-of-domain images. Inspired by the chain-of-thou…

cs.CV2025

CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos

Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang

Reasoning Video Object Segmentation is a challenging task, aiming at generating a mask sequence from an input video given a complex and implicit text query. While existing works fi…

cs.CV2025

StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams

Yang LI, Jinglu Wang, Lei Chu +4

The advent of 3D Gaussian Splatting (3DGS) has advanced 3D scene reconstruction and novel view synthesis. With the growing interest of interactive applications that need immediate…

cs.CV2025

Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts

Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang

Reasoning segmentation is a challenging vision-language task that aims to output the segmentation mask with respect to a complex, implicit, and even non-visual query text. Previous…

cs.CV2025

Beyond and Free from Diffusion: Invertible Guided Consistency Training

Chia-Hong Hsu, Shiu-hong Kao, Randall Balestriero

Guidance in image generation steers models towards higher-quality or more targeted outputs, typically achieved in Diffusion Models (DMs) via Classifier-free Guidance (CFG). However…

cs.CV2025

UVRM: A Scalable 3D Reconstruction Model from Unposed Videos

Shiu-hong Kao, Xiao Li, Jinglu Wang +4

Large Reconstruction Models (LRMs) have recently become a popular method for creating 3D foundational models. Training 3D reconstruction models with 2D visual data traditionally re…