1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2026
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
q-bio.QM2026★ 1 cited
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
Tianyu Liu, Weihao Xuan, Hao Wu +15
Advances in AI have introduced several strong models in computational pathology to usher it into the era of multi-modal diagnosis, analysis, and interpretation. However, the curren…
cs.HC2026
VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
Shunyu Liu, Minghao Liu, Huichi Zhou +31
Recent advances have showcased the extraordinary capabilities of Large Language Model (LLM) agents in tackling web-based information-seeking tasks. However, existing efforts mainly…