8 papers
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
Chenlong Deng, Mengjie Deng, Junjie Wu +10
Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This paradigm overlooks the rich dep…
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
Jihong Wang, Jiamu Zhou, Weiming Zhang +7
With the advancement of vision-language models, web automation has made significant progress. However, deploying autonomous agents in real-world settings remains challenging, prima…
OSCAR: Optimization-Steered Agentic Planning for Composed Image Retrieval
Teng Wang, Rong Shan, Jianghao Lin +8
Composed image retrieval (CIR) requires complex reasoning over heterogeneous visual and textual constraints. Existing approaches largely fall into two paradigms: unified embedding…
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
SHengjie Ma, Chenlong Deng, Jiaxin Mao +5
While reinforcement learning (RL) enhances their ability to plan and reason across retrieval steps, we identify a critical failure mode in this setting: Tool-Call Hacking. Unlike e…
Epitome: Pioneering an Experimental Platform for AI-Social Science Integration
Jingjing Qu, Kejia Hu, Jun Zhu +9
Large Language Models (LLMs) enable unprecedented social science experimentation by creating controlled hybrid human-AI environments. We introduce Epitome (www.epitome-ai.com), an…
LightAgent: Production-level Open-source Agentic AI Framework
Weige Cai, Tong Zhu, Jinyi Niu +6
With the rapid advancement of large language models (LLMs), Multi-agent Systems (MAS) have achieved significant progress in various application scenarios. However, substantial chal…