4 papers
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation
Jun Ma, Zhenye Yang, Ruichen Zhou +3
Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and senso…
STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
Zhenye Yang, Jinpeng Chen, Huan Li +6
Conversational recommender systems (CRSs) aim to proactively capture user preferences through natural language dialogue and recommend high-quality items. To achieve this, CRS gathe…
Hierarchical Intent-guided Optimization with Pluggable LLM-Driven Semantics for Session-based Recommendation
Jinpeng Chen, Jianxiang He, Huan Li +5
Session-based Recommendation (SBR) aims to predict the next item a user will likely engage with, using their interaction sequence within an anonymous session. Existing SBR models o…