3 papers
cs.AI2026
Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models
Gwang Gook Lee, Kenan Emir Ak, Jay Mohta +2
Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful tran…
cs.LG2026
Routing-Based Continual Learning for Multimodal Large Language Models
Jay Mohta, Kenan Emir Ak, Gwang Lee +3
Multimodal Large Language Models (MLLMs) struggle with continual learning, often suffering from catastrophic forgetting when adapting to sequential tasks. We introduce a routing-ba…
cs.CL2025
Don't Just Demo, Teach Me the Principles: A Principle-Based Multi-Agent Prompting Strategy for Text Classification
Peipei Wei, Dimitris Dimitriadis, Yan Xu +1
We present PRINCIPLE-BASED PROMPTING, a simple but effective multi-agent prompting strategy for text classification. It first asks multiple LLM agents to independently generate can…