4 papers
PokeFusion Attention: A Lightweight Cross-Attention Mechanism for Style-Conditioned Image Generation
Jingbang Tang
Style-conditioned text-to-image (T2I) generation with diffusion models requires both stable character structure and consistent, fine-grained style expression across diverse prompts…
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
Zeteng Lin, Xingxing Li, Wen You +4
Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this…
Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans
Hongbin Huang, Junwei Li, Tianxin Xie +8
High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present…
Emission-GPT: A domain-specific language model agent for knowledge retrieval, emission inventory and data analysis
Jiashu Ye, Tong Wu, Weiwen Chen +11
Improving air quality and addressing climate change relies on accurate understanding and analysis of air pollutant and greenhouse gas emissions. However, emission-related knowledge…