collaborators

5 papers

cs.CV2026

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

Pablo Messina, Andrés Villa, Juan León Alcázar +5

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign…

cs.CV2026

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs

Wayner Barrios, Andrés Villa, Juan León Alcázar +2

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs)…

cs.CV2025

Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Andrés Villa, Juan Carlos León Alcázar, Alvaro Soto +1

Large Vision and Language Models have enabled significant advances in fully supervised and zero-shot visual tasks. These large architectures serve as the baseline to what is curren…

cs.LG2025

Towards Faster and More Compact Foundation Models for Molecular Property Prediction

Yasir Ghunaim, Andrés Villa, Gergo Ignacz +3

Advancements in machine learning for molecular property prediction have improved accuracy but at the expense of higher computational cost and longer training times. Recently, the J…

cs.CV2025

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Andrés Villa, Juan León Alcázar, Motasem Alfarra +3

Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstream tasks. The fusion of these mod…