1 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh +4
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit…