8 papers
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive…
SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task
Lang Mei, Xiaohan Yu, Chong Chen +27
Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training eff…
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Hao Jiang, Gangtao Xin, Yingdi Huang +35
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…
MSearcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning
Xiaohan Yu, Chao Feng, Lang Mei +1
Recent advances in DeepResearch-style agents have demonstrated strong capabilities in autonomous information acquisition and synthesize from real-world web environments. However, e…
CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning
Xiaohan Yu, Zhihan Yang, Chong Chen
Multimodal Retrieval Augmented Generation (MRAG) systems have shown promise in enhancing the generation capabilities of multimodal large language models (MLLMs). However, existing…
Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation
Xiaohan Yu, Li Zhang, Xin Zhao +1
The recent breakthrough of large language models (LLMs) in natural language processing has sparked exploration in recommendation systems, however, their limited domain-specific kno…