activity
20242026
most citedCooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models

1 citations · 3 across the 43 of their papers we have counts for

collaborators

46 papers

cs.CV2026

Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding

Yizhou Liu, Fei Tang, Yuchen Yan +8

Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tu…

cs.CL2026

When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation

Yiwen Qiu, Linjuan Wu, Dingming Li +7

Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expr…

cs.CL2026

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Yuhan Wang, Zhengxi Lu, Yuchen Yan +6

Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks p…

cs.CL2026

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu, Jianze Wang +8

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large l…

cs.CL2026

BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes

Fei Tang, Huawen Shen, Zhiqiong Lu +7

Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high…

cs.AI2026

Agent-G: Gaussian Guidance for Agentic Reinforcement Learning

Zixuan Wang, Yanrui Miao, Zhengxi Lu +6

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy expl…