3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.AI2026
Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models
Gwang Gook Lee, Kenan Emir Ak, Jay Mohta +2
Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful tran…
cs.LG2025
Routing-Based Continual Learning for Multimodal Large Language Models
Jay Mohta, Kenan Emir Ak, Gwang Lee +3
Multimodal Large Language Models (MLLMs) struggle with continual learning, often suffering from catastrophic forgetting when adapting to sequential tasks. We introduce a routing-ba…
cs.CV2023★ 3 cited
NICE: CVPR 2023 Challenge on Zero-shot Image Captioning
Taehoon Kim, Pyunghwan Ahn, Sangyun Kim +39
In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project and share the results and outcomes of 2023 challenge. This project is designed t…