7 citations · 8 across the 3 of their papers we have counts for
3 papers
stat.ML2026
DDO-RM: Distribution-Level Policy Improvement after Reward Learning
Tiantian Zhang, Jierui Zuo, Michael Chen +1
Recent theory suggests that reward-model-first methods can be more sample-efficient than direct policy fitting when the reward function is statistically simpler than the induced po…
cs.LG2023★ 1 cited
Multi-Modal Machine Learning for Assessing Gaming Skills in Online Streaming: A Case Study with CS:GO
Longxiang Zhang, Wenping Wang
Online streaming is an emerging market that address much attention. Assessing gaming skills from videos is an important task for streaming service providers to discover talented ga…
cs.IR2023★ 7 cited
Integrity and Junkiness Failure Handling for Embedding-based Retrieval: A Case Study in Social Network Search
Wenping Wang, Yunxi Guo, Chiyao Shen +4
Embedding based retrieval has seen its usage in a variety of search applications like e-commerce, social networking search etc. While the approach has demonstrated its efficacy in…