7 papers
Bridging Short Videos and Live Streams: Reasoning-Guided Multimodal LLMs for Cross-Domain Representation Learning
Le Zhang, Xiaolan Zhu, Yuchen Wang +9
As live streaming services grow, many platforms offer short videos and live streams to meet diverse needs. Short videos carry substantial traffic and rich behavior signals, whereas…
LLMs Need Encoders for Semantic IDs Too
Xiangyi Chen, Zelun Wang, Xinyi Li +3
Multimodal LLMs use dedicated encoders to bridge non-language modalities (vision encoders for images, depth models for audio codec tokens) because raw token embeddings alone cannot…
SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking
Ruochen Yang, Yueyang Liu, Zijie Zhuang +14
Large-scale live-streaming recommendation requires precise modeling of non-stationary content semantics under strict real-time serving constraints. In industrial deployment, two co…
QARM V2: Quantitative Alignment Multi-Modal Recommendation for Reasoning User Sequence Modeling
Tian Xia, Jiaqi Zhang, Yueyang Liu +25
With the evolution of large language models (LLMs), there is growing interest in leveraging their rich semantic understanding to enhance industrial recommendation systems (RecSys).…
Foresight Prediction Enhanced Live-Streaming Recommendation
Jiangxia Cao, Ruochen Yang, Xiang Chen +8
Live-streaming, as an emerging media enabling real-time interaction between authors and users, has attracted significant attention. Unlike the stable playback time of traditional T…
Multimodal Recommendation via Self-Corrective Preference Alignmen
Yalong Guan, Xiang Chen, Mingyang Wang +7
With the rapid growth of live streaming platforms, personalized recommendation systems have become pivotal in improving user experience and driving platform revenue. The dynamic an…