1 citations · 1 across the 2 of their papers we have counts for
4 papers
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback
Zhen-Yu Zhang, Yuting Tang, Jiandong Zhang +2
Online reinforcement learning from human feedback (RLHF) has emerged as a promising paradigm for aligning large language models (LLMs) by continuously collecting new preference fee…
Decoupled Multimodal Fusion for User Interest Modeling in Click-Through Rate Prediction
Alin Fan, Hanqing Li, Sihan Lu +2
Modern industrial recommendation systems improve recommendation performance by integrating multimodal representations from pre-trained models into ID-based Click-Through Rate (CTR)…
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
Jiangyuan Wang, Kejun Xiao, Qi Sun +4
Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as…
In-context Demonstration Matters: On Prompt Optimization for Pseudo-Supervision Refinement
Zhen-Yu Zhang, Jiandong Zhang, Huaxiu Yao +2
Large language models (LLMs) have achieved great success across diverse tasks, and fine-tuning is sometimes needed to further enhance generation quality. Most existing methods rely…