5 papers
Latent Implicit Visual Reasoning
Kelvin Li, Chuyi Shang, Leonid Karlinsky +3
While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are l…
PuYun-LDM: A Latent Diffusion Model for High-Resolution Ensemble Weather Forecasts
Lianjun Wu, Shengchen Zhu, Yuxuan Liu +7
Latent diffusion models (LDMs) suffer from limited diffusability in high-resolution (<=0.25°) ensemble weather forecasting, where diffusability characterizes how easily a latent d…
EVLF-FM: Explainable Vision Language Foundation Model for Medicine
Yang Bai, Haoran Cheng, Yang Zhou +40
Despite the promise of foundation models in medical AI, current systems remain limited - they are modality-specific and lack transparent reasoning processes, hindering clinical ado…
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items
Minjie Zou, Sahana Srinivasan, Thaddaeus Wai Soon Lo +13
Recent advances in reasoning-focused large language models (LLMs) mark a shift from general LLMs toward models designed for complex decision-making, a crucial aspect in medicine. H…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…