6 papers
iLRM: An Iterative Large 3D Reconstruction Model
Gyeongjin Kang, Seungtae Nam, Seungkwon Yang +4
Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explicit 3D representations, such as 3…
LLMs Can Compensate for Deficiencies in Visual Representations
Sho Takishita, Jay Gala, Abdelrahman Mohamed +2
Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We inv…
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
Abdelrahman Mohamed, Yova Kementchedjhieva
Despite significant advances in vision-language models (VLMs), image captioning often suffers from a lack of detail, with base models producing short, generic captions. This limita…
JEEM: Vision-Language Understanding in Four Arabic Dialects
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali +7
We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Mo…
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy +24
In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divi…
fCOP: Focal Length Estimation from Category-level Object Priors
Xinyue Zhang, Jiaqi Yang, Xiangting Meng +2
In the realm of computer vision, the perception and reconstruction of the 3D world through vision signals heavily rely on camera intrinsic parameters, which have long been a subjec…