2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 2 cited
Chain-of-Thought Reasoning is a Policy Improvement Operator
Hugh Zhang, David C. Parkes
Large language models have astounded the world with fascinating new capabilities. However, they currently lack the ability to teach themselves new skills, relying instead on large…
cs.GT2022★ 1 cited
A Simple Adaptive Procedure Converging to Forgiving Correlated Equilibria
Hugh Zhang
Simple adaptive procedures that converge to correlated equilibria are known to exist for normal form games (Hart and Mas-Colell 2000), but no such analogue exists for extensive-for…