From the 1 of 11 linked papers with an AI index.
11 papers
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models
Chonghan Qin, Xiachong Feng, Ziyun Song +3
Large language models (LLMs) are routinely prompted to take on social roles ranging from individuals to institutions, yet it remains unclear whether their internal representations…
SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution
Xiachong Feng, Yi Jiang, Xiaocheng Feng +9
Social intelligence, the ability to navigate complex interpersonal interactions, presents a fundamental challenge for language agents. Training such agents via reinforcement learni…
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
Xiachong Feng, Deyi Yin, Xiaocheng Feng +9
Games offer a compelling paradigm for developing general reasoning capabilities in language models, as they naturally demand strategic planning, probabilistic inference, and adapti…
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
Qiushi Sun, Zhoumianze Liu, Chang Ma +18
Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdisciplinary research. Recently, various…
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models
Chonghan Qin, Xiachong Feng, Weitao Ma +2
Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. T…