3 papers
cs.CV2026
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
Jiaqi Deng, Zonghan Wu, Huan Huo +1
Knowledge-based Vision Question Answering (KB-VQA) extends general Vision Question Answering (VQA) by not only requiring the understanding of visual and textual inputs but also ext…
cs.MA2026
MetaForge: A Self-Evolving Multimodal Agent that Retrieves, Adapts, and Forges Tools On Demand
Shouang Wei, Houcheng Min, Xinpeng Dong +8
Multimodal agents have achieved notable progress on complex reasoning tasks through tool use, yet remain limited by two issues: statically predefined tool inventories fail to gener…
cs.IR2025
Cloud-Device Collaborative Agents for Sequential Recommendation
Jing Long, Sirui Huang, Huan Huo +3
Recent advances in large language models (LLMs) have enabled agent-based recommendation systems with strong semantic understanding and flexible reasoning capabilities. While LLM-ba…