activity
20232025
most citedKaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

3 citations · 5 across the 8 of their papers we have counts for

collaborators

7 papers

cs.LG2025

Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards

Zetian Sun, Dongfang Li, Zhuoen Chen +2

Reward sparsity in long-horizon reinforcement learning (RL) tasks remains a significant challenge, while existing outcome-based reward shaping struggles to define meaningful immedi…

cs.AI2025

Improving Value-based Process Verifier via Low-Cost Variance Reduction

Zetian Sun, Dongfang Li, Baotian Hu +1

Large language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, rem…

cs.AI2025

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

Zetian Sun, Dongfang Li, Xuhui Chen +2

The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize t…

cs.LG2025

Improving Value-based Process Verifier via Structural Prior Injection

Zetian Sun, Dongfang Li, Baotian Hu +2

In the Large Language Model(LLM) reasoning scenario, people often estimate state value via Monte Carlo sampling. Though Monte Carlo estimation is an elegant method with less induct…

cs.CL20253 cited

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

Xinshuo Hu, Zifei Shan, Xinping Zhao +10

As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, pri…

cs.CL2024

CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models

Dongfang Li, Zetian Sun, Xinshuo Hu +2

Large Language Models (LLMs) need to adapt to the continuous changes in data, tasks, and user preferences. Due to their massive size and the high costs associated with training, LL…