6 papers
Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models
Gwang Gook Lee, Kenan Emir Ak, Jay Mohta +2
Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful tran…
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13
Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusabl…
PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
Tuo Zhang, Alin-Ionut Popa, Yan Xu +2
Large language model (LLM)-based agents frequently generate seemingly coherent plans that fail upon execution due to infeasible actions, constraint violations, and compounding erro…
Routing-Based Continual Learning for Multimodal Large Language Models
Jay Mohta, Kenan Emir Ak, Gwang Lee +3
Multimodal Large Language Models (MLLMs) struggle with continual learning, often suffering from catastrophic forgetting when adapting to sequential tasks. We introduce a routing-ba…
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
Peizhou Huang, Zixuan Zhong, Zhongwei Wan +12
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA…
Leveraging Uncertainty Estimation for Efficient LLM Routing
Tuo Zhang, Asal Mehradfar, Dimitrios Dimitriadis +1
Deploying large language models (LLMs) in edge-cloud environments requires an efficient routing strategy to balance cost and response quality. Traditional approaches prioritize eit…