3 papers
cs.CV2026
Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
TsaiChing Ni, ZhenQi Chen, YuanFu Yang
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufactur…
cs.CV2025
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
ZhenQi Chen, TsaiChing Ni, YuanFu Yang
Recent text-to-image diffusion models have achieved remarkable visual fidelity but often struggle with semantic alignment to complex prompts. We introduce CritiFusion, a novel infe…
cs.GR2025
DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
Zhen-Qi Chen, Yuan-Fu Yang
With the rapid advancement of diffusion-based generative models, Stable Diffusion (SD) has emerged as a state-of-the-art framework for high-fidelity im-age synthesis. However, exis…