◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zhi Zhang

4 papers hereh-index 2142 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.CL1
  • cs.LG1
same name
  • Zhi Zhang — 18 papers, h 18
  • Zhi Zhang — 11 papers, h 11
  • Zhi Zhang — 9 papers, h 5
  • Zhi Zhang — 8 papers, h 1
  • Zhi Zhang — 8 papers, h 3
  • Zhi Zhang — 7 papers, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedSeed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

1 citations · 1 across the 3 of their papers we have counts for

collaborators

4 papers

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Aili Chen, Aonian Li, Baichuan Zhou +215

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

cs.LG2026

Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning

Zhi Zhang, Zhen Han, Costas Mavromatis +9

Reinforcement learning (RL) plays a central role in large language model (LLM) post-training. Among existing approaches, Group Relative Policy Optimization (GRPO) is widely used, e…

cs.CL2025★ 1 cited

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

ByteDance Seed, :, Jiaze Chen +267

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 8…

cs.AI2024

Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization

Ruijie Xu, Zhihan Liu, Yongfei Liu +4

We address the challenge of online Reinforcement Learning from Human Feedback (RLHF) with a focus on self-rewarding alignment methods. In online RLHF, obtaining feedback requires i…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.