4 papers
PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
Tuo Zhang, Alin-Ionut Popa, Yan Xu +2
Large language model (LLM)-based agents frequently generate seemingly coherent plans that fail upon execution due to infeasible actions, constraint violations, and compounding erro…
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
Peizhou Huang, Zixuan Zhong, Zhongwei Wan +12
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA…
Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
Ziming Dai, Tuo Zhang, Fei Gao +5
The growing industrial demand for customized and cost-efficient large language models (LLMs) is fueled by the rise of vertical, domain-specific tasks and the need to optimize perfo…
Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
Tuo Zhang, Yuechun Sun, Ruiliang Liu
In this work, we present a retrieval-augmented generation (RAG)-based system for provenance analysis of archaeological artifacts, designed to support expert reasoning by integratin…