4 papers
SceneFoundry: Generating Interactive Infinite 3D Worlds
ChunTeng Chen, YiChen Hsu, YiWen Liu +5
The ability to automatically generate large-scale, interactive, and physically realistic 3D environments is crucial for advancing robotic learning and embodied intelligence. Howeve…
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
YuChe Hsu, AnJui Wang, TsaiChing Ni +1
We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language pro…
Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
TsaiChing Ni, ZhenQi Chen, YuanFu Yang
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufactur…
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
ZhenQi Chen, TsaiChing Ni, YuanFu Yang
Recent text-to-image diffusion models have achieved remarkable visual fidelity but often struggle with semantic alignment to complex prompts. We introduce CritiFusion, a novel infe…