3 papers
cs.CV2026
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
Qingtao Pan, Kai Ye, Zhihao Dou +2
In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image synthesis. However, such cons…
cs.CV2026
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
Qingtao Pan, Zhihao Dou, Shuo Li
Large Multimodal Models (LMMs) struggle to adapt varying computational budgets due to numerous visual tokens. Previous methods attempted to reduce the number of visual tokens befor…
cs.CV2025
Geometric Origins of Bias in Deep Neural Networks: A Human Visual System Perspective
Yanbiao Ma, Bowei Liu, Andi Zhang
Bias formation in deep neural networks (DNNs) remains a critical yet poorly understood challenge, influencing both fairness and reliability in artificial intelligence systems. Insp…