◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Qi Zhang

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author2

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.LG1
  • cs.MA1
  • cs.RO1
ORCID 0000-0003-2519-5580
same name
  • Qi Zhang — 46 papers
  • Qi Zhang — 25 papers, h 24
  • Qi Zhang — 13 papers
  • Qi Zhang — 10 papers
  • Qi Zhang — 9 papers
  • Qi Zhang — 8 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20222025
most citedCommunication-Efficient Actor-Critic Methods for Homogeneous Markov Games

2 citations · 4 across the 4 of their papers we have counts for

collaborators

4 papers

cs.LG2025

BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping

Zhiheng Xi, Xin Guo, Yang Nan +18

Reinforcement learning (RL) has recently become the core paradigm for aligning and strengthening large language models (LLMs). Yet, applying RL in off-policy settings--where stale…

cs.RO2025

A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning

Shaopeng Zhai, Qi Zhang, Tianyi Zhang +7

Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLA…

cs.CL2023★ 2 cited

Universal Multi-modal Entity Alignment via Iteratively Fusing Modality Similarity Paths

Bolin Zhu, Xiaoze Liu, Xin Mao +4

The objective of Entity Alignment (EA) is to identify equivalent entity pairs from multiple Knowledge Graphs (KGs) and create a more comprehensive and unified KG. The majority of E…

cs.MA2022★ 2 cited

Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games

Dingyang Chen, Yile Li, Qi Zhang

Recent success in cooperative multi-agent reinforcement learning (MARL) relies on centralized training and policy sharing. Centralized training eliminates the issue of non-stationa…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.