6 papers
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation
Yuxuan Yao, Yuxuan Chen, Hui Li +6
Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visua…
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
Yuwei Sun, Yuxuan Yao, Hui Li +1
Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following tasks remains constrained. Whi…
Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering
Wuzhenghong Wen, Bowen Zhou, Jinwen Huang +5
Question Answering over Temporal Knowledge Graphs (TKGQA) has attracted growing interest for handling time-sensitive queries. However, existing methods still struggle with: 1) weak…
Reinforcement Learning Enhanced Multi-hop Reasoning for Temporal Knowledge Question Answering
Wuzhenghong Wen, Chao Xue, Su Pan +2
Temporal knowledge graph question answering (TKGQA) involves multi-hop reasoning over temporally constrained entity relationships in the knowledge graph to answer a given question.…
MCM: Multi-layer Concept Map for Efficient Concept Learning from Masked Images
Yuwei Sun, Lu Mi, Ippei Fujisawa +4
Masking strategies commonly employed in natural language processing are still underexplored in vision tasks such as concept learning, where conventional methods typically rely on f…
Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task
Wuzhenghong Wen, Su Pan, yuwei Sun
Schema linking is a critical step in Text-to-SQL task, aiming to accurately predict the table names and column names required for the SQL query based on the given question. However…