Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
cs.CV2024
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
Qilong Zhangli, Jindong Jiang, Di Liu +6
While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substanti…