3 papers
cs.CL2026
An Empirical Study of VLM Pipelines for Long-Document QA
Kenan E. Ak, Jay Mohta, Gwang Gook Lee +2
Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them me…
cs.AI2026
Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models
Gwang Gook Lee, Kenan Emir Ak, Jay Mohta +2
Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful tran…
cs.LG2025
Routing-Based Continual Learning for Multimodal Large Language Models
Jay Mohta, Kenan Emir Ak, Gwang Lee +3
Multimodal Large Language Models (MLLMs) struggle with continual learning, often suffering from catastrophic forgetting when adapting to sequential tasks. We introduce a routing-ba…