collaborators

6 papers

cs.AI2026

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

Gwang Gook Lee, Kenan Emir Ak, Jay Mohta +2

Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful tran…

cs.AI2026

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13

Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusabl…

cs.AI2026

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement

Tuo Zhang, Alin-Ionut Popa, Yan Xu +2

Large language model (LLM)-based agents frequently generate seemingly coherent plans that fail upon execution due to infeasible actions, constraint violations, and compounding erro…

cs.LG2026

Routing-Based Continual Learning for Multimodal Large Language Models

Jay Mohta, Kenan Emir Ak, Gwang Lee +3

Multimodal Large Language Models (MLLMs) struggle with continual learning, often suffering from catastrophic forgetting when adapting to sequential tasks. We introduce a routing-ba…

cs.CV2026

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

Peizhou Huang, Zixuan Zhong, Zhongwei Wan +12

Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA…

cs.NI2025

Leveraging Uncertainty Estimation for Efficient LLM Routing

Tuo Zhang, Asal Mehradfar, Dimitrios Dimitriadis +1

Deploying large language models (LLMs) in edge-cloud environments requires an efficient routing strategy to balance cost and response quality. Traditional approaches prioritize eit…