activity
20222026
most citedUniParser: A Unified Log Parser for Heterogeneous Log Data

124 citations · 169 across the 83 of their papers we have counts for

collaborators
Showing cs.LGShow all

17 papers · 1 filter

cs.LG2026

ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory

Yue Fang, Zhibang Yang, Fangkai Yang +5

Large language model (LLM) agents increasingly rely on external tools served by shared providers and accessed by heterogeneous downstream agents. Existing approaches improve tool u…

cs.LG2026

Beyond State Consistency: Behavior Consistency in Text-Based World Models

Youling Huang, Guanqiao Chen, Junchi Yao +8

World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and offline evaluation. In text-…

cs.LG2025

Towards Active Synthetic Data Generation for Finetuning Language Models

Samuel Kessler, Menglin Xia, Daniel Madrigal Diaz +5

A common and effective means for improving language model capabilities involves finetuning a ``student'' language model's parameters on generations from a more proficient ``teacher…

cs.LG2025

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs

Qibin Wang, Pu Zhao, Shaohan Huang +6

Test-time scaling (TTS) has gained widespread attention for enhancing LLM reasoning. Existing approaches such as Best-of-N and majority voting are limited as their performance depe…

cs.LG2025

VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model

Jiani Zheng, Lu Wang, Fangkai Yang +7

Training Vision-Language Models (VLMs) for Graphical User Interfaces (GUI) agents via Reinforcement Learning (RL) faces critical challenges: environment-based RL requires costly in…

cs.LG2025

Pretrain Value, Not Reward: Decoupled Value Policy Optimization

Chenghua Huang, Lu Wang, Fangkai Yang +6

In this paper, we explore how directly pretraining a value model simplifies and stabilizes reinforcement learning from human feedback (RLHF). In reinforcement learning, value estim…