24 papers
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
Chenlong Deng, Mengjie Deng, Junjie Wu +10
Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This paradigm overlooks the rich dep…
SpecTran: Spectral-Aware Transformer-based Adapter for LLM-Enhanced Sequential Recommendation
Yu Cui, Feng Liu, Zhaoxiang Wang +4
Traditional sequential recommendation (SR) models learn low-dimensional item ID embeddings from user-item interactions, often overlooking textual information such as item titles or…
Discrete Preference Learning for Personalized Multimodal Generation
Yuting Zhang, Ying Sun, Dazhong Shen +6
The emergence of generative models enables the creation of texts and images tailored to users' preferences. Existing personalized generative models have two critical limitations: l…
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
Chenyu Zhou, Huacan Chai, Wenteng Chen +18
Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the…
OThink-SRR1: Search, Refine and Reasoning with Reinforced Learning for Large Language Models
Haijian Liang, Zenghao Niu, Junjie Wu +3
Retrieval-Augmented Generation (RAG) expands the knowledge of Large Language Models (LLMs), yet current static retrieval methods struggle with complex, multi-hop problems. While re…
Sharpness-Aware Minimization for Generalized Embedding Learning in Federated Recommendation
Fengyuan Yu, Xiaohua Feng, Yuyuan Li +3
Federated recommender systems enable collaborative model training while keeping user interaction data local and sharing only essential model parameters, thereby mitigating privacy…