3 papers
cs.AI2026
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
Jianing Wang, Jianfei Zhang, Qi Guo +24
We introduce LongCat-Flash-Prover, a flagship 560-billion-parameter open-source Mixture-of- Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-…
cs.CL2025
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
Jianfei Zhang, Bei Li, Jun Bai +4
In-Context Learning (ICL) empowers Large Language Models (LLMs) for rapid task adaptation without Fine-Tuning (FT), but its reliance on demonstration selection remains a critical c…
cs.CL2024
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
Jianfei Zhang, Jun Bai, Bei Li +4
Aligning Large Language Models (LLMs) with general human preferences has been proved crucial in improving the interaction quality between LLMs and human. However, human values are…