3 papers
cs.CV2026
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
cs.AI2024
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…
cs.CV2024
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
Qilong Zhangli, Jindong Jiang, Di Liu +6
While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substanti…