most citedBlack-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AI2026

PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor

Qianjun Pan, Junyi Wang, Jie Zhou +10

To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key…

cs.AI20252 cited

Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories

Shilian Chen, Jie Zhou, Tianyu Huai +9

Model merging refers to the process of integrating multiple distinct models into a unified model that preserves and combines the strengths and capabilities of the individual models…

cs.AI2025

Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark

Yuxuan Cai, Yipeng Hao, Jie Zhou +14

As AI advances toward general intelligence, the focus is shifting from systems optimized for static tasks to creating open-ended agents that learn continuously. In this paper, we i…

cs.LG2025

Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback

Yutao Yang, Jie Zhou, Junsong Li +5

This paper introduces an interactive continual learning paradigm where AI models dynamically learn new skills from real-time human feedback while retaining prior knowledge. This pa…

cs.AI2025

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law

Qianjun Pan, Wenkai Ji, Yuyang Ding +8

This survey explores recent advancements in reasoning large language models (LLMs) designed to mimic "slow thinking" - a reasoning process inspired by human cognition, as described…

cs.CL2025

Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning

Junsong Li, Jie Zhou, Yutao Yang +7

Automatic math correction aims to check students' solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answ…