activity
20242026
most citedKaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

3 citations · 3 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL2026

LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding

Gang Lin, Dongfang Li, Zhuoen Chen +4

The proliferation of long-context large language models (LLMs) exposes a key bottleneck: the rapidly expanding key-value cache during decoding, which imposes heavy memory and laten…

cs.AI2025

Improving Value-based Process Verifier via Low-Cost Variance Reduction

Zetian Sun, Dongfang Li, Baotian Hu +1

Large language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, rem…

cs.AI2025

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

Zetian Sun, Dongfang Li, Xuhui Chen +2

The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize t…

cs.CL2025

Take Off the Training Wheels Progressive In-Context Learning for Effective Alignment

Zhenyu Liu, Dongfang Li, Xinshuo Hu +4

Recent studies have explored the working mechanisms of In-Context Learning (ICL). However, they mainly focus on classification and simple generation tasks, limiting their broader a…

cs.LG2025

Improving Value-based Process Verifier via Structural Prior Injection

Zetian Sun, Dongfang Li, Baotian Hu +2

In the Large Language Model(LLM) reasoning scenario, people often estimate state value via Monte Carlo sampling. Though Monte Carlo estimation is an elegant method with less induct…

cs.CL20253 cited

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

Xinshuo Hu, Zifei Shan, Xinping Zhao +10

As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, pri…