2 citations · 2 across the 8 of their papers we have counts for
4 papers · 1 filter
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
Yibo Wang, Guangda Huzhang, Yuwei Hu +7
Recent advances in Multimodal Large Language Models (MLLMs) have substantially driven the progress of autonomous agents for Graphical User Interface (GUI). Nevertheless, in real-wo…
Establishing Reliability Metrics for Reward Models in Large Language Models
Yizhou Chen, Yawen Liu, Xuesi Wang +5
The reward model (RM) that represents human preferences plays a crucial role in optimizing the outputs of large language models (LLMs), e.g., through reinforcement learning from hu…
Exploit Customer Life-time Value with Memoryless Experiments
Zizhao Zhang, Yifei Zhao, Guangda Huzhang
As a measure of the long-term contribution produced by customers in a service or product relationship, life-time value, or LTV, can more comprehensively find the optimal strategy f…
Imitate TheWorld: A Search Engine Simulation Platform
Yongqing Gao, Guangda Huzhang, Weijie Shen +4
Recent E-commerce applications benefit from the growth of deep learning techniques. However, we notice that many works attempt to maximize business objectives by closely matching o…