6 papers
Efficient Causal Structure Learning via Modular Subgraph Integration
Haixiang Sun, Pengchao Tian, Zihan Zhou +3
Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face c…
Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification
Xixian Wu, Yang Ou, Pengchao Tian +4
Vision-language models (VLMs) have demonstrated significant potential in Visual Question Answering (VQA). However, the susceptibility of VLMs to hallucinations can lead to overconf…
MeshRipple: Structured Autoregressive Generation of Artist-Meshes
Junkai Lin, Hang Long, Huipeng Guo +8
Meshes serve as a primary representation for 3D assets. Autoregressive mesh generators serialize faces into sequences and train on truncated segments with sliding-window inference…
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
Weihang Wang, Xinhao Li, Ziyue Wang +5
Object hallucination in Large Vision-Language Models (LVLMs) significantly impedes their real-world applicability. As the primary component for accurately interpreting visual infor…
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Yu Xie, Jielei Zhang, Pengyu Chen +5
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to…
MX-Font++: Mixture of Heterogeneous Aggregation Experts for Few-shot Font Generation
Weihang Wang, Duolin Sun, Jielei Zhang +1
Few-shot Font Generation (FFG) aims to create new font libraries using limited reference glyphs, with crucial applications in digital accessibility and equity for low-resource lang…