8 papers
Image2Gcode: Image-to-G-code Generation for Additive Manufacturing Using Diffusion-Transformer Model
Ziyue Wang, Yayati Jadhav, Peter Pak +1
Mechanical design and manufacturing workflows conventionally begin with conceptual design, followed by the creation of a computer-aided design (CAD) model and fabrication through m…
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Yu Xie, Jielei Zhang, Pengyu Chen +5
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to…
3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis
Ziyue Wang, Linghan Cai, Chang Han Low +8
3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling…
When Noise Lowers The Loss: Rethinking Likelihood-Based Evaluation in Music Large Language Models
Xiaosha Li, Chun Liu, Ziyu Wang
The rise of music large language models (LLMs) demands robust methods of evaluating output quality, especially in distinguishing high-quality compositions from "garbage music". Cur…
Video-KTR: Reinforcing Video Reasoning via Key Token Attribution
Ziyue Wang, Sheng Jin, Zhongrong Zuo +5
Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequen…
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
Weihang Wang, Xinhao Li, Ziyue Wang +5
Object hallucination in Large Vision-Language Models (LVLMs) significantly impedes their real-world applicability. As the primary component for accurately interpreting visual infor…